Language models can do more than generate answers from patterns learned during training. Meta’s Toolformer is designed to decide for itself when an outside tool could help, which tool to call, and what information to send. The idea is to make tool use part of the model’s learned behavior rather than confining it to a narrow set of tasks.
Learning when a tool can help
Connecting tools to language models is not new. The distinctive feature of Toolformer is that it learns to use a range of tools with limited human-written examples, then selects when a call is useful for a particular text task.
The training process begins with a language model learning a “handful” of human-written API statements. It uses these examples to label a larger dataset with possible tool actions. Researchers then automatically select useful examples from that pool to fine-tune Toolformer.
The process used about 25,000 examples per API. This gives the model practice in choosing a tool and supplying its arguments, while also learning to incorporate the returned result into its next-token predictions.
A toolkit for different tasks
After training, the researchers reported that the model could call several kinds of services: a calculator, a question-answering system, two search engines including a Wikipedia search, a translation system, and a calendar.
Tool use depends on the task. The model decides whether a call is needed and when to make it, rather than calling an API for every prompt. In principle, that can help with tasks where a dedicated tool can provide information or computation that the language model would otherwise have to produce from its learned patterns alone.
The researchers tested whether this approach could improve performance without task-specific examples at inference time, a setting described as zero-shot. They reported that a GPT-J model with 6.7 billion parameters, when using tools, outperformed GPT-3 with 175 billion parameters on selected tasks.
The reported gains were not universal. Tool-assisted improvements became apparent in tests starting at about 775 million parameters, while smaller models performed similarly with and without tools. The Wikipedia API was an exception for question-answering tasks, which the researchers suggested could reflect how easy that API was to use. Even so, Toolformer did not beat GPT-3 on the QA benchmarks.
What the approach could address
External services could help with areas where language models have difficulty, including reliable math problem-solving and fact-checking. A calculator can handle a calculation, for example, while a search or question-answering service can provide information for a response. Toolformer’s contribution is a method for learning when such help may be relevant.
The paper also reports that tool use improved zero-shot performance, and that the difference between predictions made with and without API calls remained substantial even for the largest model tested. This suggests that access to tools and model size can contribute in different ways: growing models improve at tasks on their own, while their ability to make use of APIs also improves.
Limits on independent tool use
Toolformer’s calls are not yet flexible enough to support every workflow. It cannot use one tool’s output as the input to a second tool, because instructions for each API are generated independently. That rules out tasks that require a series of linked operations.
It also cannot interact with a tool’s results and revise its next action. For instance, it cannot inspect a search engine’s many results and then refine the query based on what it found. Tool use is also sensitive to the exact wording of a prompt, so a change in phrasing may affect whether the model calls a tool.
Efficiency is another limitation. Processing more than a million documents produced only a few thousand meaningful examples of calculator API calls. The system also does not account for the computational cost of making an API call when deciding whether to use one.
Taken together, the results show both the promise and the boundaries of self-directed tool use. A model can learn to choose from several APIs and gain on selected tasks, but the reported system still makes isolated calls rather than carrying out an interactive sequence of steps.