Some questions about an image cannot be answered by looking at its pixels alone. Identifying an object may be only the first step; answering questions about it can require finding facts elsewhere. Google’s AVIS is a system designed to search for that missing visual information by combining computer vision with web and image search.
When an image needs outside information
Large language models have helped make image captioning and visual question answering possible. Yet visual language models can struggle with questions that require both interpreting a real-world image and retrieving knowledge that is not readily available in it. Researchers call this challenge “visual information seeking.”
Examples include questions about when an airline was founded or the year a car was built. Those details are not necessarily visible in an image. A system must connect what it can identify visually to relevant information elsewhere, then decide whether it has enough evidence to answer.
AVIS approaches this as a sequence of decisions. It uses computer vision tools to extract information from an image, web search to retrieve facts, and image search to look for information in metadata associated with visually similar images.
A system that adjusts its search
AVIS brings together Google’s PALM language model and those search and vision tools. Instead of following a fixed, two-step process, it can plan and reason iteratively, adapting its next action in response to what it has learned from previous tool use.
The framework has three parts: a planner, a working memory and a reasoner. The planner selects the next action, including an API call and query. Working memory retains information from earlier API executions, while the reasoner processes those outputs to identify useful information.
In each cycle, the reasoner updates the system’s understanding and the planner uses that state to choose a tool and query. The process continues until the reasoner judges that enough information has been gathered to provide an answer. This gives AVIS a way to change course as it searches, rather than committing to a single search plan in advance.
Learning from how people use tools
To guide the system’s behavior, the researchers studied how people make decisions while using visual reasoning tools. They used patterns in people’s actions to construct a transition graph, which guides AVIS through possible tool-use sequences.
That design connects tool choice to the current state of the search. A result from one step can inform what the system tries next, while working memory preserves information gathered along the way. Together, the planner and reasoner help organize a search that may involve visual analysis as well as outside sources.
Results and remaining limits
On the Infoseek dataset, AVIS achieved 50.7% accuracy, outperforming fine-tuned visual language models such as OFA and PaLI. On the OK-VQA dataset, it achieved 60.2% accuracy with few examples, outperforming most previous work and approaching fine-tuned models, Google said.
These results show how combining a language model with external tools can help with visual questions that call for information beyond an image. They do not remove the practical cost of the approach: the PALM model used by AVIS has 540 billion parameters and is computationally intensive.
The team wants to explore the framework on other reasoning tasks and investigate whether lighter language models could perform these capabilities. That work would address a key question raised by AVIS: how much of its tool-guided reasoning can be retained when the underlying model requires fewer computational resources?