Large language models are earning serious respect in mathematics, but leading mathematicians are drawing a clear boundary around what they can do. Timothy Gowers and Peter Sarnak both credit LLMs with meaningful mathematical ability, while arguing that the hardest kind of creative reasoning remains out of reach.
What LLMs are good at in mathematics
The strongest case for LLMs in math is not that they think like great mathematicians. It is that they can work across known material, combine established methods, and explore many possible directions.
Gowers describes current models as effective at putting together techniques that already exist. They can also try many search paths, which matters in mathematics because progress often depends on testing different routes through a problem.
That makes LLMs useful in familiar problem spaces. When a task can be addressed by drawing on existing theory, a model may be able to assemble a workable route, derive results, or help navigate a complex body of prior knowledge.
This is still a substantial capability. Mathematics contains many problems where the challenge is not inventing an entirely new foundation, but finding the right combination of known tools.
Where the limit appears
The concern raised by Gowers is that a vast search space is not enough on its own. A model may be able to try many possibilities, but the central challenge is knowing which few paths are worth pursuing.
That choice depends on intuition. In the view described in the source article, current LLMs do not have the kind of intuition needed to identify the productive route before spending effort on countless unpromising ones.
Sarnak’s criticism points in a similar direction. He accepts that AI can derive results from existing theory. The harder test is different: starting from an elementary question and developing the abstractions that later support major proofs.
That distinction matters because major mathematical advances are not only about calculation or derivation. They often require a new way to frame the problem. Once that framing exists, many steps may become expressible through known logic. Before it exists, the work is much less mechanical.
The problem of new assumptions
DeepMind researcher Tom Zahavy reaches a related conclusion in the paper "LLMs Can't Jump." The bottleneck he identifies is "manipulative abduction," described in the source as the ability to invent new foundational assumptions with no linguistic precedent.
That phrase captures the difference between extending a pattern and creating a new starting point. LLMs are trained on language, so they are naturally strong where useful precedents exist in language. But a mathematical breakthrough may require assumptions that are not already available as a recognizable linguistic pattern.
The source article notes that world models could offer a possible path forward. It does not present them as a solved answer, but as one direction that might address the gap between manipulating existing language and forming new foundational views of a problem.
This is why the debate is not simply about whether LLMs can do math. The sharper question is what kind of math they can do, and what kind of mathematical work still depends on capacities that current systems do not clearly possess.
Why benchmarks do not settle the question
These assessments feed into a broader argument about AI progress. One side of the debate asks whether LLMs are becoming genuinely more versatile. The other asks whether they are "just" getting better at benchmarks and familiar problem spaces.
That distinction is especially important in mathematics because benchmark success can reward the ability to recognize known patterns, combine familiar techniques, and search quickly. Those are valuable skills, but they are not the same as inventing the abstractions behind major proofs.
For users of AI tools, the practical lesson is straightforward. LLMs may be strong assistants when the work involves existing methods, derivations, and exploration of known approaches. They are less reliable as engines for the kind of original conceptual leap that defines the deepest mathematical creativity.
The result is a more measured picture of AI in mathematics. Large language models are not being dismissed as simple calculators. They are being treated as powerful systems with real mathematical strengths and a still-unsolved weakness: creating the new ideas that make some proofs possible in the first place.