AI Pushes the Riemann Hypothesis Further Without Solving It

An unreleased Anthropic model made reported progress on the Riemann hypothesis, a famous unsolved problem about prime numbers. It did not produce a general proof, but it increased the lower bound of solutions for which the hypothesis holds true and renewed debate over AI’s role in mathematics.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story shows AI gaining useful problem-solving capability in mathematics, but with no clear danger or societal degradation beyond mild dependence concerns.

AI Pushes the Riemann Hypothesis Further Without Solving It

An unreleased Anthropic model has made reported progress on the Riemann hypothesis, one of mathematics’ most durable open problems. The result does not settle the hypothesis, and the $1 million bounty for a working general proof remains unclaimed. But the way the work was produced points to a larger question now facing mathematics: how should the field treat discoveries made with large language models?

What the Anthropic model achieved

The Riemann hypothesis has resisted proof for more than 150 years. At its core, it concerns the distribution of prime numbers, making it one of the central unsolved questions in mathematics.

Anthropic said on Monday that an as-yet-unreleased model had made significant progress on the problem. According to the source article, the model significantly increased the lower bound of solutions for which the Riemann hypothesis holds true.

That distinction matters. The model did not deliver the full general proof that would resolve the hypothesis. Instead, it pushed verified progress further along a specific line: showing that the hypothesis holds for more solutions than previously established by this reported work.

For a problem with this history, that is still notable. The Riemann hypothesis is not a puzzle where partial progress is easy to dismiss. Even limited advances can become important because they clarify what has been tested, what remains open, and how future attempts might be structured.

How the work was done

The process described by Anthropic is almost as important as the mathematical result. An Anthropic staff member without significant mathematical training prompted the model to "take a real stab" at proving the hypothesis. The model then coordinated the work across the following day and a half.

In that period, the system tested 650 different ideas for solving the problem. It coordinated across 60 sub-agents and spent 31 million in total.

A footnote to the paper describes how the work was divided among those agents. Two subagents developed the key mathematical ideas. Another 13 contributed ideas to those agents. 30 tried but could not develop new ideas. 13 acted as validators to check the correctness of the arguments. The final two helped write the initial paper.

This division of labor is significant because it shows a model being used less like a single chatbot and more like a managed research process. Some parts generated candidate approaches. Others checked them. Others helped produce the written result.

The finding was confirmed by two of Anthropic’s in-house mathematicians. It was also formalized using Lean, the open-source proof assistant. That gives the reported result an additional layer of mathematical structure beyond a conversational exchange with an AI system.

Why this fits a wider AI mathematics trend

The Riemann hypothesis result is not presented as an isolated event. It follows a series of recent mathematical results involving large language models, or LLMs.

The source article notes that a number of Erdos problems have been solved by AI models over the course of this year. It also says that stronger model releases have led to more impressive results.

OpenAI recently released a set of ten major results proved by its internal "Astra" model. Separately, an Anthropic effort disproved the long-standing Jacobian conjecture.

Taken together, these examples point to a shift in what contemporary AI systems can contribute. The claim is not that LLMs now solve every difficult mathematical problem. The more careful point is that they can explore ideas, organize attempts, and sometimes produce results that specialists consider worth checking and formalizing.

That is why the Riemann hypothesis case is likely to draw attention beyond the specific lower-bound improvement. It suggests that AI may be useful not only for routine assistance, but also for work closer to discovery.

The credit and responsibility problem

The growing role of AI in mathematics is creating both excitement and concern. The source article points to a public declaration signed in June by prominent mathematicians. The declaration warned that AI could weaken critical values in the field.

The central concern is attribution and responsibility. Mathematics has traditionally attached proofs to authors who take credit for the discovery and accept responsibility for the correctness of the argument. AI-assisted work complicates that model, especially when a system coordinates many sub-agents and produces ideas that humans later inspect.

In this case, the prompt came from an Anthropic staff member without significant mathematical training. The model organized the exploration. Two in-house mathematicians confirmed the result. Lean was used to formalize it. Each of those roles matters, but they do not fit neatly into the familiar picture of an individual mathematician proving a theorem.

Not everyone sees that as a purely negative development. In a blog post responding to the declaration, Fields Medal winner Timothy Gowers questioned whether AI might reshape mathematics in a more complex and positive way.

"If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all," Gowers wrote.

What this means now

The immediate takeaway is simple: the Riemann hypothesis remains unsolved. The $1 million bounty for a working general proof remains unclaimed, and contemporary AI models still cannot close the problem.

But the Anthropic result shows that the boundary is moving. An unreleased model, guided by a relatively simple prompt, reportedly explored hundreds of approaches, coordinated dozens of sub-agents, and produced work that mathematicians at Anthropic checked and formalized.

That does not end the debate over AI in mathematics. It sharpens it. If LLMs can contribute meaningful mathematical ideas without fitting the field’s traditional model of authorship, mathematicians will have to decide how to evaluate, credit, and trust that work.

For now, the Riemann hypothesis remains what it has been for more than 150 years: a major unsolved problem. The difference is that AI systems are no longer only watching from the sidelines. They are beginning to take part in the search.