Artificial intelligence is pushing mathematics toward a question that cannot be answered by another benchmark: what is the field trying to achieve?
In a new essay for the 2026 International Congress of Mathematicians, Terence Tao argues that the math community should move beyond the narrow debate over AI capability. The deeper issue, in his view, is whether established mathematical values can survive a world where machines may produce more research outputs than people can comfortably evaluate.
A crisis about values, not truth
Tao compares the moment to the foundational crisis that shook mathematics between 1900 and 1930. Russell's paradox and Gödel's incompleteness theorems forced mathematicians to make explicit assumptions that had often remained unstated. The result was a rigorous framework that has lasted for a century.
The challenge Tao sees now is different. It is not mainly about whether mathematical truth exists or how it can be formalized. It is about the working culture around mathematics: what counts as a contribution, what deserves reward, what it means to understand a result, and whether a machine can be credited with doing mathematical work.
His working hypothesis is direct: "AI tools will, reasonably soon, become capable of performing a reasonable fraction of research-level mathematical tasks, with reasonable levels of success, quality, supervision, and cost." If that proves right, the field will face a practical and philosophical test at the same time.
Early evidence from AI proof attempts
Tao points to the First-Proof Project as evidence that this shift is not merely theoretical. In the second round, ten never-published research problems were tested against four AI systems under controlled conditions.
The result was notable: Seven of the ten received at least one passing grade from at least one system. In this context, that meant a solution considered essentially flawless or needing only minor revisions. The reported costs were in the tens to hundreds of dollars per problem.
That does not settle the future of AI in mathematics. It does, however, make the question harder to ignore. If systems can increasingly generate acceptable answers to research problems, then the old scarcity of proofs may no longer organize the field in the same way.
When mathematical metrics become targets
Tao's concern is that mathematics has always pursued several goals at once. Solving problems, building theories, sustaining a research community, and training younger mathematicians have typically reinforced one another. AI could separate those goals.
That is where Goodhart's law enters his argument: "When a measure becomes a target, it ceases to be a good measure." In mathematics, a solved problem or polished proof can serve as a visible signal of progress. But Tao warns that generative AI is especially suited to optimizing for the appearance of such signals.
The incentive structure around AI makes this more difficult. The industry rewards benchmarkable and quotable achievements. Those are exactly the kinds of outputs that can look like clean evidence of mathematical progress while leaving deeper goals, such as understanding and community judgment, harder to measure.
Proof abundance could create a new bottleneck
If AI-generated proofs become common, mathematics may move from a shortage of proofs to an excess of them. Tao warns that proofs could accumulate faster than experts can check, read, or absorb.
The Erdős problem database already contains dozens of AI-generated submissions that no human expert has volunteered to verify. That example points to a simple operational problem: a proof has limited value if the community lacks the time or willingness to evaluate it.
There is also a subtler issue with proofs polished by AI. Tao observes that human-written proofs often reveal where the difficult work happened. A careful lemma, a change of notation, or a heavily revised paragraph can help readers locate the real mathematical pressure.
An AI-polished proof may remove those cues. It can become "easy to read and hard to learn from." Tao's point is not that messy writing is inherently better. It is that certain rough edges can show the reader how the argument was built. As he puts it, the "mistakes" in human exposition "can be genuinely helpful to the reader."
Why explanation may become a publication standard
For practical guidance, Tao points to the Leiden Declaration on Artificial Intelligence and Mathematics, published in June 2026 and backed by the International Mathematical Union. The declaration appears in his argument as part of a broader push to define responsible norms before AI-generated mathematics becomes routine.
Tao's own rule of thumb is stricter than simply asking whether a proof checks out. "If the authors cannot convincingly demonstrate that they are able to give a clear, expert-level talk on their results, one that is correct and properly attributed, then the result should not be published." On that view, a proof that no human can properly explain remains incomplete, even if it has been formally verified.
Training is another pressure point. Tao argues that young mathematicians need special protection for the "irreducibly human aspect" of mathematical work, with AI tool use kept tightly restricted. Producing correct homework is not the same as becoming a mathematician.
At the same time, Tao does not present AI as something to reject entirely. He discloses using AI for literature search, diagram creation, text completion, and converting his slides into paper format. The distinction is between using tools to support mathematical work and allowing tool outputs to replace the human understanding that gives the work meaning.
That is why Tao's essay moves beyond a simple capability debate. Timothy Gowers and Peter Sarnak have credited large language models with real mathematical abilities while also seeing limits around genuinely new ideas, a pattern that appears in other research as well. Tao's central question is broader: before AI changes the supply of proofs, mathematics may need to say more clearly what it wants proofs to do.