OpenAI has published hundreds of mathematical proofs produced by its models, but the release has renewed a basic question: who checks that the explanations and their formal versions say the same thing? Recommendations from a group of prominent mathematicians call for human understanding and careful review alongside AI-generated results.
Guidelines set expectations for AI math
The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton University’s Institute for Advanced Studies, includes nine researchers at institutions around the world. It released guidelines for frontier labs working on mathematical problems at the end of September.
In a statement about OpenAI’s latest proofs, AGMAI said the mathematical community would ultimately assess whether its recommendations were followed. The group’s first request was to stop testing advanced mathematical problems on proprietary models. OpenAI’s release says it is evaluating its proprietary models on open research problems in mathematics.
The lab did follow some of the group’s principles. It released results promptly and included information about how the models reached their conclusions. But the release did not meet every recommendation described in the article.
Transparency and formalization remain uneven
OpenAI released 719 manuscripts, but only 10 included the model’s chain of thought. That leaves readers with limited access to the reasoning behind most of the results, even as the lab provides some information about how its models reached their conclusions.
AGMAI also recommended formalizing proofs that people do not understand. Formalization expresses a mathematical argument in a system that can check it mechanically. In this release, 42% of the proofs had not undergone that process.
The group’s recommendations also called for machine-readable metadata that connects natural-language explanations with formal artifacts. OpenAI did not include that metadata in these releases. Such links matter because readers need to compare the explanation with the version submitted to a proof-checking system.
A translation can introduce questions
AI systems may first produce a natural-language explanation and then translate it into Lean, a programming language that can confirm a proof’s accuracy by compiling it as code. A paper released this week by mathematicians at the University of Cambridge and King’s College in London examines whether that translation can preserve the original argument.
The paper describes at least two discrepancies between OpenAI’s natural-language proof and the Lean code for a problem derived from the Navier-Stokes equations, which describe complex fluid behavior. The discrepancies do not necessarily disprove either solution. They do raise questions about relying on models to formalize their own work without human involvement.
The authors of the “lost in translation” paper say OpenAI’s natural-language proof and other automatically formalized Lean proofs should not be trusted without peer review and scrutiny comparable to that given to other proofs. Their concern is that a proof’s explanation and its coded form may not match exactly.
Mathematical results need people to carry them forward
For mathematicians, a result involves more than presenting a possible solution. Human researchers take responsibility for new findings and discuss them with the broader community through papers, talks and seminars. That process can deepen understanding, identify approaches useful for other problems and help apply new knowledge in practical fields.
Terence Tao, a prominent mathematician who has criticized OpenAI’s approach, wrote on social media that some people prompting AI to solve problems have little interest in the wider field after a target is declared solved. He also said they may not understand the output well enough to answer questions, give talks or engage with other researchers.
Harvard University mathematics professor Melanie Wood told TechCrunch that when a model produces a solution to a hard problem, “there is not human understanding of them at the point of release, and now the work begins.” That gap puts the focus on what happens after publication: researchers need time and support to examine the results and explain what they mean.
AGMAI suggested that OpenAI help fund human mathematicians’ work so the lab’s solutions can become meaningful to the field. The group did not provide TechCrunch with a fuller evaluation of this proof release. For now, the central issue remains whether the release process supports the human understanding and review that mathematical results require.