Language models are becoming stronger at solving formal problems and extracting patterns from data. But Tom Zahavy of Google Deepmind argues that this is not enough to produce a scientific revolution.
In a position paper titled "LLMs can't jump", Zahavy says the missing ingredient is not more fluency or better deduction. It is the ability to invent a new explanatory frame when no ready-made concept exists in language yet.
The missing step in discovery
Zahavy builds his argument around a view of discovery that Albert Einstein described in a letter to Maurice Solovine. In that account, scientific thought begins with sensory experience, moves through an intuitive leap toward axioms, and then uses logic to derive conclusions that can be tested.
Axioms are the basic assumptions on which a theory rests. They are not proved inside the theory itself. The hard part, in Zahavy's framing, is not always applying logic after the axioms are known. It is finding the right axioms in the first place.
To clarify the gap, Zahavy uses a distinction from Charles Sanders Peirce between three kinds of reasoning: deduction, induction, and abduction.
- Deduction starts from fixed rules and derives conclusions from them.
- Induction observes patterns in examples and generalizes from them.
- Abduction proposes an explanation for something surprising.
Language models are well suited to induction because they are trained to detect statistical structure. Zahavy also argues that formal derivation is increasingly within reach. Systems such as AlphaProof, Gemini, and GPT-5 now achieve gold-level scores on International Mathematical Olympiad problems.
That still leaves the most difficult kind of abduction. Ordinary abduction selects a likely explanation from options that are already available, as a doctor might match symptoms to a disease. Zahavy says language models can do that. The deeper problem is what he calls "manipulative abduction": creating a cause or concept for which there is not yet a linguistic template.
Why optimization may not be enough
Zahavy uses Einstein's work on general relativity to show why this matters. He even concedes that a language model could probably derive general relativity if Einstein's assumptions were already supplied. But that would not solve the central problem of discovery. The machine would still be starting after the decisive conceptual move had already happened.
AI systems often improve by comparing a prediction with reality, measuring the error, and adjusting. This works when there is a clear gap between output and target. Zahavy argues that Einstein's situation did not provide that kind of signal.
Newton's physics had been confirmed with extreme precision. The known anomaly was a tiny shift in Mercury's orbit, and that had been attributed to a hypothetical hidden planet called "Vulcan". In Zahavy's view, an optimization-driven system would have little reason to reject the existing framework. It would likely search for a smaller fix inside that framework.
The evidence that later supported Einstein's theory, including Eddington's measurement of light deflection, came years after the theory had been formulated. That timing is central to Zahavy's claim. A breakthrough can arrive before the confirming data makes the old model visibly untenable.
Why the body matters
Zahavy points to Einstein's "happiest thought" as the kind of event language models lack. Einstein imagined a freely falling observer who no longer feels gravity. He also imagined a physicist inside an accelerating elevator in space, leading to the conclusion that acceleration and gravity are indistinguishable from the inside.
For Zahavy, this was not merely a calculation. It was an embodied simulation: a mental replay of physical experience that helped generate a new foundation for thought.
He draws a similar line to Archimedes. In the familiar story, Archimedes does not reach his buoyancy principle by calculation alone. The principle emerges from the physical experience of water rising as he steps into a bathtub.
The common point is that both cases involve more than rearranging symbols. They involve sensory grounding. Zahavy compares language models to John Searle's "Chinese Room", where a person follows rules for handling Chinese characters without understanding the language. In his argument, language models can manipulate the symbols of physics without sharing the physical experience that gives those symbols meaning.
Where current AI systems fall short
Zahavy does not dismiss scientific automation. Sakana's AI Scientist and Deepmind's AlphaEvolve are presented as impressive systems. But he draws a boundary around what they can do.
According to the source argument, the AI Scientist recombines existing concepts. AlphaEvolve optimizes effectively, but it still depends on a clear error signal that can be reduced step by step. Einstein, in Zahavy's telling, did not have that signal when the key conceptual work was done.
This distinction matters because scientific revolutions are not just faster searches through existing ideas. They can require a new way to describe the problem itself. If a system depends on language templates or measurable error, it may remain trapped inside the assumptions it is supposed to overturn.
World models as the possible route forward
Zahavy sees physically consistent world models as a possible path beyond this limitation. He separates them from video generators such as Veo, which predict likely next frames. In that kind of system, a falling apple falls because similar continuations are common in training data, not because the model understands gravity.
The more promising direction, in his view, is action-controllable world models such as Genie. These systems let an agent intervene in a simulation and run counterfactual experiments. Zahavy's example is mentally cutting an elevator cable.
That kind of synthetic lab could give AI a route toward the feedback needed for manipulative abduction. Instead of only predicting text or extending visual patterns, an agent could act, observe consequences, and test imagined alternatives inside a physically consistent environment.
Zahavy's argument does not say that language models are useless for science. It says their strongest abilities may sit on the wrong side of the hardest step. They can recognize patterns, work through rules, and automate parts of research. But the leap that creates a new framework may require models grounded in simulated physical interaction, not language alone.