Can Mathematicians Verify OpenAI’s Flood of New Results?

OpenAI released nearly 400 AI-generated mathematical results across 719 manuscripts, spanning many fields. Researchers say the scale, uneven formal verification and difficulty of assessing the papers could leave the work under review for years.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 2 ►

The release’s scale and uneven verification strain human review, mildly raising concerns about dependence on AI-generated work and the effort needed to assess it.

Can Mathematicians Verify OpenAI’s Flood of New Results?

OpenAI’s latest mathematics release is large enough to pose a problem of its own: researchers first have to work out what it contains. The collection includes nearly 400 AI-generated results spread across 719 manuscripts, and mathematicians say evaluating the work will take time, especially when proofs have not been formally checked.

A collection too large to assess at a glance

The manuscripts cover fields including combinatorics, geometry, number theory, theoretical computer science, algebra, topology, probability and statistical mechanics, and mathematical physics. OpenAI published guidance for navigating the sprawling GitHub repository, but even getting an overview is a substantial task.

Some mathematicians told The Verge that reading the roughly 40-page contents list and abstracts took much of an hour. Álvaro Lozano-Robledo, a professor of mathematics at the University of Connecticut, described reviewing the abstracts as overwhelming. That initial scan is only the start: understanding a result requires examining its argument and how it relates to existing work.

The volume also stretches across areas of mathematics that require different kinds of expertise. Researchers must decide which claims are relevant to their fields, whether the reasoning holds, and how the findings should be understood. Taken together, those tasks help explain why mathematicians expect the collection to take years to make sense of.

Formal proofs can help, but still need review

Some manuscripts include formalizations in Lean, a programming language and proof assistant that can let results be checked computationally. A formal proof can give researchers confidence that a claim follows logically, even when the full argument is difficult for them to follow.

But the release is not uniformly formalized. OpenAI said the results were at different stages of verification and that many, but not all, manuscripts had been formalized. It reported that 300 top-line results out of 719 manuscripts had been formalized, around 42 percent, and said it would add more formalizations as they were obtained.

Even a Lean formalization does not settle every question. Researchers still need to check that the formal proof establishes the same claim made in the accompanying manuscript. Mathematicians who examined the material told The Verge that the quality varied, and that some verified statements did not appear to match the papers’ claims neatly.

That leaves researchers with several steps to assess: understand the paper, inspect any formal proof and check that both describe the same result. Where no formalization is available, readers must evaluate the mathematical argument directly. Either way, verification takes time.

Quantity raises concerns about quality

Kevin Buzzard, a mathematics professor at Imperial College London, said that among the algebraic number theory results he identified, only around six immediately stood out to him. Few, if any, appeared to have Lean verification. He said that without formal checks, he would need to read work that might be wrong, wait for others to assess it, or wait for someone to formalize it before confirming whether it was correct.

Other mathematicians also raised concerns about AI-generated material that is hard to read, confusing or inaccurate. They said such work can show little understanding of its subject and can be especially poor at crediting other researchers. Some had used the term “slopocalypse” for the anticipated flood of papers.

Initial reactions to this release were reportedly better than some researchers expected, though that comparison came against a backdrop of criticism of OpenAI’s earlier mathematical write-ups, including concerns about attribution. A better first impression does not establish that every paper is clear, correct or up to the standards expected for significant academic claims.

Review takes people, expertise and time

When formal verification is missing, a paper becomes the main way for mathematicians to check, understand and put a result in context. That means the writing itself matters: readers need to be able to follow the argument and see what is being claimed. Several researchers described papers they found difficult to follow, including work in areas they knew well.

The challenge is also one of capacity. The models can produce mathematical work quickly across many specialties, while OpenAI’s human staff do not have the breadth of expertise or resources to scrutinize every finding at the frontier of mathematics, researchers told The Verge. The resulting workload then falls in part to outside mathematicians, who must choose what to examine and how to check it.

For the field, the immediate question is not just how many results the system produced. It is how those claims will be verified, understood and credited, and whether that review can keep pace if more work arrives before the current collection has been assessed. Researchers interviewed by The Verge said that even if the AI systems disappeared, understanding the release could occupy mathematicians for years.