Meta’s Nougat Turns Research PDFs Into Searchable Text

Meta’s Nougat model converts images of scientific papers into structured, machine-readable text. It handles formulas and tables better than a cited alternative, though the team says consistency and repetitive output remain challenges.

WTF Index IDIOCRACY
◄ Terminator 0 Idiocracy 1 ►

Nougat automates making research papers searchable, with no clear harm and only a mild risk of greater dependence on automated processing.

Meta’s Nougat Turns Research PDFs Into Searchable Text

Scientific papers often arrive as PDFs designed for people to read, which can make their contents harder for software to process. Meta’s Nougat model aims to close that gap by turning page images into structured text, including elements such as formulas and tables.

Reading a scientific page as a whole

Nougat stands for Neural Optical Understanding for Academic Documents. It applies optical character recognition (OCR) to scientific documents, using a variant of Vision Transformer for image analysis.

Rather than reading a page one line at a time, Nougat processes the full page. The research team says this approach helps it interpret details in mathematical notation, including superscripts and subscripts that traditional OCR has often transcribed incorrectly.

The result is intended to be more than plain text copied from a page. Nougat converts PDF images into structured, machine-readable text, so software can work with the content of research articles more directly.

Training on papers and their source code

To train the model, the team used scientific-article PDFs from sources such as arXiv and PubMed Central, paired with the corresponding LaTeX source code from the authors. The training dataset contains more than 8 million pages.

That pairing gives the model examples of how the visual page relates to the underlying document text. It is especially relevant for scientific papers, where a page can combine prose with mathematical expressions, tables, and formatting that carry meaning.

Making those features machine-readable could help improve access to scientific knowledge. When papers can be processed as structured text, their contents may be easier for other tools and workflows to use than when they remain embedded in page images.

Strong results, with uneven performance

In tests described by the team, Nougat extracted continuous text with a BLEU score of over 91% and accuracy of over 96%. Its results for formulas and tables were lower, at just over 75%.

Even in those more difficult categories, the reported performance was substantially higher than that of GROBID for mathematical formulas. The source reports GROBID’s accuracy for formulas at just under 11%.

These figures point to a useful distinction: Nougat performed most strongly on continuous text, while formulas and tables remained harder to reproduce. The comparison suggests progress in handling scientific layouts, but the lower scores for structured elements show that extraction is not equally reliable across every kind of content.

What still needs work

The researchers also identify problems that remain. One is maintaining consistency across documents. Another is avoiding repetitive text loops during generation, in which output can repeat instead of progressing through the page content.

Those challenges matter because reliable conversion depends on more than recognizing individual words. A useful machine-readable version also needs to preserve the relationships and structure that make a scientific paper understandable.

Meta presents Nougat as a way to make millions of scientific articles more accessible by bridging the divide between PDF pages and structured text. The code and models are available on GitHub, and the project page provides more information and examples. The model’s reported results offer a basis for further work in scientific document processing, while its stated limitations show where that work can continue.