Modern AI can write, code, summarize, and solve many problems with striking fluency. But fluency is not the same as reasoning. The central lesson from AlphaGo is that powerful machine intelligence may need more than bigger models and longer answers: it may need a way to test possibilities, track uncertainty, and revise beliefs in a form people can inspect.
What AlphaGo Did That Looked Impossible
In March 2016, AlphaGo played Lee Sedol in Seoul in a five-game Go match. During game two, the program placed a stone on the fifth line of the board in a move that appeared so strange that some commentators wondered whether the system had malfunctioned.
That move, remembered as move 37, was not an error. AlphaGo won the game and went on to defeat Lee Sedol 4-1. Lee, one of the greatest professional Go players of all time, later said: “I thought AlphaGo was based on probability calculation and that it was merely a machine,” and added, “But when I saw this move, I changed my mind. Surely, AlphaGo is creative.”
The comparison with chess matters. When Deep Blue beat then reigning world chess champion Garry Kasparov in 1997, it searched six to eight moves ahead per player and evaluated 200 million chess positions per second using human-coded rules. Go is much harder to brute-force because the value of a stone depends on distant groups, territory, and long sequences of future play. Searching even a fraction of all possible outcomes would require a supercomputer billions of years.
AlphaGo therefore had to do something different. It needed to judge the board quickly while also looking beyond the moves that seemed normal to human experts.
The Split Between Intuition and Search
AlphaGo combined two systems. Its policy network learned to estimate what move a strong human player might make. That part of the system treated move 37 as highly unlikely, with a roughly one in 10,000 chance of being played by an expert human.
The move was chosen because another part of AlphaGo looked ahead. Its search machinery explored a game tree with thousands of branches, each branch representing a possible future. Instead of accepting the most human-looking choice, AlphaGo weighed what could happen after candidate moves and countermoves.
This resembles a behavioral-science distinction popularized by Daniel Kahneman. System 1 thinking is fast, intuitive, and effortless. System 2 thinking is slower, more deliberate, and step-by-step. AlphaGo had a machine version of that division: neural networks supplied candidate judgments, while search tested those judgments against possible futures.
Neither side was enough by itself. Intuition alone would not have selected move 37. Search alone would have faced too many possible moves. The strength came from the combination.
Why Today’s LLMs Are Different
Large language models work in another way. They generate the next token again and again. That makes them excellent at pattern completion across subjects that people write about, but it also makes them closer to System 1 than to true deliberation.
After ChatGPT appeared, it became clear that language fluency alone was not enough for many useful tasks. One response was chain of thought: models generate intermediate steps before giving a final answer. Those steps can break a problem into pieces, carry partial results forward, and shape what comes next.
The gains have been real, especially in mathematics and coding. But the source article argues that this still does not create a separate reasoning system. The intermediate steps are produced by the same next-token prediction process, just continued for longer before the model answers.
Three limits are especially important:
- No persistent inspectable state. The model does not usually keep an explicit ledger of hypotheses, confidence levels, evidence, and open questions that can be systematically updated.
- No clean split between knowledge and manipulation. What the system knows and how it uses that knowledge are tangled inside neural-network weights.
- Unreliable explanations. Chatbots can produce chains of thought that look like deliberation, while research has shown that they may concoct them after the fact.
This distinction matters because a convincing explanation is not the same as a traceable reasoning process.
Why Auditability Matters in High-Stakes AI
In fields such as medicine, engineering, and scientific research, the route to an answer can matter as much as the answer itself. If a system makes a mistake in medical diagnosis and treatment, users need to know what failed. Did it reason incorrectly? Did it rely on invalid evidence? Did it start from the wrong assumptions?
The article argues that future AI systems need an explicit epistemic state. That means a representation of what the system treats as settled, what it doubts, what it has rejected, and which questions remain open. Reasoning would then become a sequence of moves that changes that state: deducing consequences, dividing problems into parts, and deciding what question, calculation, or experiment should come next.
AlphaGo offers a model for this idea. Its game tree recorded the possible futures it had examined, with moves and positions annotated by judgments from its neural networks. As its reasoning progressed, the tree changed, and the final move came from synthesizing what that structure contained.
Real-world reasoning is harder than Go or chess. Outside a board game, the current state is only partly known, possible actions are broad and variable, and consequences may be stochastic or unknown. Still, the article argues that recent advances in LLMs and other neural models make a more general reasoning architecture possible.
What a More Trustworthy System Would Need
LLMs could still play an important role in such systems. They can suggest ways to approach a problem using known information and available resources. They can work with tools through APIs or code. They can also help assess whether available evidence supports a claim.
But another part of the system would need to evaluate each reasoning move independently. It would ask whether the move truly reduces uncertainty and would update beliefs only when evidence supports the change. Under those rules, a system could accumulate certified knowledge and improve by learning from earlier reasoning experiences.
The argument is not that bigger System 1 models are useless. Scale can sharpen intuition. But making intuition larger does not automatically make it deliberate. Move 37 mattered because a machine maintained a position, considered possible futures, and chose a move its own instinctive component would likely have dismissed.
That is the deeper challenge for AI in drug discovery, materials, climate, and diagnosis. The systems needed for those domains must do more than produce persuasive language. They must reach conclusions through an auditable sequence of evidence, inference, and belief revision.