External Search Helps LLMs Find Better Paths to Solutions

Microsoft researchers and collaborators developed Everything of Thought (XOT), a framework that pairs a language model with an external reasoning module inspired by AlphaZero. In tests on three puzzles, XOT outperformed other approaches, though it was not perfectly reliable.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

XOT improves puzzle-solving with external search, but the routine research result shows no clear lean toward danger or human dependency.

External Search Helps LLMs Find Better Paths to Solutions

Large language models can be prompted to break a hard problem into smaller steps. A method called Everything of Thought (XOT) takes a different approach: it gives the language model help from an external module that searches for useful reasoning structures. Researchers report that the approach performed better than other methods on several puzzle tasks, while still falling short of perfect reliability.

Moving the search outside the language model

Some prompting techniques ask a model to lay out intermediate steps, or to consider multiple possible paths through a problem. XOT instead uses a separate module to search for structures of thoughts that could lead to a solution. The method was developed by researchers at Microsoft, Georgia Institute of Technology, and East China Normal University.

The external module draws on reinforcement learning and Monte Carlo Tree Search (MCTS), with an approach inspired by AlphaZero. Rather than asking the language model to generate and assess every possible line of reasoning on its own, XOT aims to give it candidate thought structures drawn from a search process.

How XOT learns to suggest solution paths

During training, MCTS explores possible solutions to a task, such as a puzzle. The search records information about thought nodes, including their states, values, and how often they are visited. That record becomes training data for reinforcement learning.

The trained model learns to predict which solution paths are likely to succeed. The intended benefit is that it can suggest promising structures on later problems without repeating an exhaustive search across the full solution tree each time. The researchers say this could help the system generalize to new problems within a game.

Once connected to a language model, the external model supplies thought structures for a problem the language model has posed. The language model can review those structures and ask for revisions, making the process collaborative. XOT therefore shifts much of the work of exploring and evaluating possible thoughts away from the language model.

Strong results, with reliability limits

The researchers evaluated XOT on the Game of 24, the 8-Puzzle, and the Pocket Cube. They report that it significantly outperformed other approaches in these tests, including on problems that those approaches could not solve.

Those results do not mean every answer is correct. XOT did not reach 100% reliability, so the tests show an improvement in the selected tasks rather than a guarantee of dependable reasoning in every setting. The source does not provide performance figures, so the size of the gains cannot be quantified here.

What the approach could mean for AI reasoning

XOT offers one way to combine a language model with external domain knowledge and search. If a separate module can find useful paths and the language model can inspect and refine them, the language model may have less work to do when solving structured problems. The researchers describe the framework as improving performance, efficiency, and flexibility together.

Whether that combination carries over beyond the tested puzzles remains an open question. The source says it is not known if or when Microsoft plans to use the method in its products. It also notes that Google DeepMind CEO Demis Hassabis said in an interview that the company would like to incorporate ideas from AlphaGo into Gemini, but it does not establish that Gemini uses XOT.