A new benchmark puts AI assistants to work on tasks that may seem ordinary to people: reasoning through a question, navigating the internet, or choosing a tool. In its first evaluation, GPT-4 struggled with these challenges, while human participants answered far more successfully.
A benchmark built around everyday capability
Researchers from Metas AI Research (FAIR), HuggingFace, AutoGPT, and GenAI introduced GAIA, short for General AI Assistants. Its premise is that a system approaching general artificial intelligence should be able to outperform people even on tasks an average person can solve.
That emphasis differs from a common direction in AI benchmarking, where systems are asked to tackle problems that are difficult for people or demand technical expertise. GAIA instead aims to examine a broader set of human-like capabilities and make results easier to interpret.
The benchmark contains 466 questions. They test basic reasoning, handling different modalities, internet navigation and tool use, including internet search. The questions are meant to challenge AI while remaining conceptually simple for humans, with fact-based answers that are concise and unambiguous.
Three levels measure the work involved
GAIA groups its questions by the number of steps and tools needed. Level 1 typically involves no tool or one tool and no more than five steps. Level 2 usually takes about five to ten steps and combines tools. Level 3 allows for sequences of actions of any length, any number of tools and general access to the world.
The researchers designed the questions for zero-shot answering, which simplifies evaluation. Because the answers have clear factual targets, assessment can be quick and objective, according to the benchmark description.
This setup tests more than whether a model can produce a plausible sentence. It also asks whether the system can work through a task, decide when a tool is useful and coordinate steps. The tiered structure makes the difference between a short question and a multi-step assignment visible.
GPT-4’s results reveal a gap
In the first evaluation, GPT-4 with plugins succeeded on only about 30 percent of the Level 1 tasks and none of the Level 3 tasks. Across all difficulty levels, GPT-4 averaged 15 percent success. Human test subjects averaged 92 percent.
Plugins are tools a language model can use when its built-in capabilities are not enough. GPT-4 performed best when the plugins were defined by people for the task, a result the researchers see as evidence of the potential in tool research.
The article points to Meta AI’s Toolformer, introduced in February 2023, as an example of work on models that learn to select tools such as calculators, question-and-answer systems or internet search. It also notes that an OpenAI study published in March on LLMs and the job market discussed language models that rely on tools.
The human comparison has a limitation: every participant had an academic background. The group included 61% with a Bachelor's, 26% with a Master's and 17% with a PhD. The article says 37.7% of people over the age of 25 in the U.S. had a bachelor's degree in 2022. Whether academic education affects performance on these benchmark tasks remains an open question.
Search and reasoning remain open challenges
The researchers also consider whether language models could serve as search engine replacements. For Level 1 questions, a person may be able to infer an answer from direct text in search results. More complex Level 2 and Level 3 queries are less suited to that approach.
Even where a person could find an answer through web search, they might take longer than an LLM assistant because they need to sift through the initial results. But speed alone does not settle the comparison: the assessment does not account for the reliability or accuracy of search results, which the article identifies as a central concern for using LLMs as search replacements.
The GAIA results suggest that performance on complex or specialized benchmarks does not by itself establish broad competence. An assistant may need to combine reasoning, tools and online information reliably to handle straightforward tasks. The researchers expect solving GAIA to mark a milestone in AI development, while the first evaluation shows a substantial gap between GPT-4 and the human participants studied.