Why LLM Test Scores Can Overstate What AI Can Do

Large language models can score highly on human-designed tests, but those results do not necessarily show how they solve problems or what they understand. Researchers argue for new evaluations that test models across varied tasks and rule out simpler explanations such as memorization.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 2 ►

The story warns that test scores can overstate AI capability and distort judgments about human-like intelligence, with a mild concern about eroding truth in evaluation.

Why LLM Test Scores Can Overstate What AI Can Do

A high score on an exam can make a large language model look capable of reasoning like a person. But the score alone cannot tell us what the model has learned, how it arrived at an answer, or whether it can handle a small change to the question. Researchers say evaluations need to probe those limits before test results are treated as evidence of broad intelligence.

A test score leaves important questions open

When psychologist Taylor Webb tried GPT-3 on abstract problems in early 2022, its performance surprised him. In later work, he and colleagues found that GPT-3 performed better than a group of undergrads on some tests of analogical reasoning, the ability to spot a relationship and use it to solve a new problem.

Those results sit alongside reports that GPT-4 performed well on academic and professional assessments, including high school tests and the bar exam. OpenAI also worked with Microsoft to show that GPT-4 could pass parts of the United States Medical Licensing Examination. Other researchers have examined abilities such as working through a problem step by step and guessing what another person believes.

These findings are striking, but human tests come with assumptions about what a score means. A person who does well may have the knowledge or skill the test was designed to measure. With a language model, a strong score might instead reflect patterns in its training data, repetition of familiar answers, or some other process. The result does not settle which explanation is right.

That uncertainty matters because test results can shape claims about whether AI systems might take on work done by teachers, journalists, and lawyers. Researchers quoted in the source argue that comparing models with people can encourage people to see human-like abilities where the evidence is less clear.

Familiar questions can reward familiarity

Large language models are trained on extensive collections of text. If questions and answers from an assessment, or similar examples, appear in that material, a model may have an advantage that has little to do with solving an unfamiliar problem. OpenAI said it checked whether tests used with GPT-4 contained text from its training data, and used paywalled questions for work on the medical exam. The source notes that these precautions cannot rule out exposure to similar questions.

One comparison involving coding tests illustrates the concern. Machine-learning engineer Horace He found that GPT-4 scored 10/10 on Codeforces tests posted before 2021 and 0/10 on tests posted after 2021. The model's training data included text collected before 2021. That contrast raised questions about whether success on older material showed memorization rather than a general ability to solve coding problems.

Webb tried to address this issue by creating new tests. In one study, his team adapted Raven’s Progressive Matrices, which ask people to find a pattern among shapes and apply it to another set. The researchers represented shape, color, and position as sequences of numbers so the tasks would not appear in training data.

But the choice of format raises another question. AI researcher Melanie Mitchell argued that replacing images with number sequences removes the visual element of the original puzzles. A model’s success on the encoded tasks therefore does not automatically show that it can solve the same kind of visual problem a person sees in Raven’s tests.

Small changes can reveal brittle performance

Another challenge is that a model can give a convincing answer to one version of a question and fail when the details change. The source describes a test in which GPT-4 was asked to stack a book, nine eggs, a laptop, a bottle, and a nail. Mitchell tried a different combination—toothpick, bowl of pudding, glass of water, and marshmallow—and got a proposed stack that was plainly unstable.

The contrast does not establish what the model can or cannot do across every physical task. It does show why a single successful answer is weak evidence of a general skill. A task can reveal performance on that particular prompt; broader claims require evidence from related tasks and variations.

Debate over theory-of-mind tests shows a similar problem. Michal Kosinski tested whether GPT-3 could distinguish what was inside a bag from what a person believed was inside, based on a misleading label. GPT-3 completed the prompts in a way Kosinski interpreted as evidence of a basic theory of mind.

Other researchers changed details of the scenario. Tomer Ullman, for example, made the bag transparent or said Sam could not read. In these variations, GPT-3 struggled to assign Sam the appropriate mental state when the situation required extra steps of reasoning. Such counterexamples make it harder to interpret success on the original version as proof of a stable ability.

Build evaluations around evidence, not resemblance

Researchers cited in the article propose making evaluations more rigorous and testing a claimed ability through several related tasks. Lucy Cheke suggests drawing on methods used to study animals: researchers run controlled experiments, examine what information is being used, and test explanations one by one. The same principle could apply to language models, even though their tasks involve language rather than navigating a maze.

Laura Weidinger and colleagues are adapting approaches used to assess cognitive abilities in preverbal human infants. The excerpt ends as it begins to describe their approach: breaking a test of an ability into a battery of related tests. The larger point is that evaluation should examine how performance changes across conditions, rather than rely on one score as a stand-in for understanding.

Human-designed exams can still show whether a model answers particular questions well. The harder task is determining what that success means. New tests, carefully varied prompts, and checks for exposure to familiar material can help make claims about large language models more precise—and make clear where the evidence remains uncertain.