Why AI grades well before it proves students learned

A Bocconi University experiment found that GPT-4o helped freshmen earn higher grades on a short business assignment. But the study did not show whether students learned more, and it exposed a gap between polished output and genuine understanding.

WTF Index IDIOCRACY
◄ Terminator 0 Idiocracy 3 ►

The story highlights AI boosting polished student output without proving learning, raising concerns about dependence and weakened genuine understanding.

Why AI grades well before it proves students learned

A classroom experiment at Bocconi University points to a problem educators can no longer treat as theoretical: AI can improve the work students submit, even when it is unclear whether it improves what they know.

The study found that GPT-4o gave students a clear advantage on a business assignment. A separate teaching intervention on causal reasoning did something different: it did not raise traditional grades, but it pushed students toward more varied ideas and more explicit thinking about why their recommendations might work.

What the Bocconi experiment tested

In November 2025, 13 sections of an introductory management course were randomly assigned to one of four conditions: a control group, a causal reasoning lesson, access to GPT-4o, or both the lesson and GPT-4o.

The participants were 1,053 freshmen. Their task was narrow and practical: write marketing recommendations for the university's merchandise shop in up to 180 words.

The causal reasoning lesson focused on the structure behind a recommendation. Students were taught to think about coherent causal logic, falsifiability, and the path from a proposed action to a desired result. In plain terms, the lesson asked students to explain not only what should be done, but why it should work and when it might not.

That distinction matters because the assignment tested more than writing. It also tested how a student turns a business problem into a short, structured proposal. GPT-4o was especially useful in producing that kind of polished answer.

GPT-4o raised scores, but the reason is complicated

Students who used GPT-4o scored nearly a full point higher on a 1-to-5 scale. Their responses included about two more ideas on average, showed stronger logical coherence, and were closer to the recommendations made by three subject-matter experts.

The researchers also looked beyond the raw score. Even after controlling for argumentation quality, number of ideas, idea diversity, and text properties, the GPT-4o advantage was still measurable. The authors attributed that remaining edge to higher content quality, not greater student knowledge.

That is the central tension. Better submitted work is not the same thing as better learning. An AI-assisted answer can be more complete, more organized, and more aligned with expert expectations while still leaving open the question of how much the student understood independently.

For schools, that makes the final paper a less reliable signal. If a grading system rewards structure, completeness, and polished reasoning in the submitted product, AI can help students perform well inside that system.

The causal reasoning lesson changed the work in another way

The lesson on causal reasoning did not improve traditional scores. In fact, work from those students scored slightly worse on average. But the lesson appeared to move students toward a different kind of answer.

Students who received the lesson more often explained why a proposed action should work and under what conditions it might fail. They also produced ideas that were more diverse and less similar to what other students wrote.

That suggests the lesson improved qualities that the ordinary score did not fully reward. The students were not simply adding more recommendations. They were more likely to make the logic behind a recommendation visible.

Combining the causal reasoning lesson with GPT-4o did not create an additional gain in traditional scores. But on causal reasoning markers, the two approaches complemented each other, and the added idea diversity from the lesson group remained.

What the grading system rewarded

The findings show a mismatch between the skills educators may value and the traits the grading rubric actually rewarded. More ideas, stronger coherence, and greater idea diversity inside a single answer were associated with higher scores.

Other qualities moved in the opposite direction. Stronger falsifiability, more detailed explanations of how actions would work, and greater divergence from other students' ideas were associated with lower scores.

That does not mean originality was always punished. The narrower point is that this assignment's traditional score mainly favored answers that were well structured and stayed within the expected solution space.

If originality, causal reasoning, and multiple approaches are supposed to matter, they have to be made explicit in grading criteria. Otherwise, students may be rewarded for producing the kind of conventional, polished response that AI can generate particularly well.

What the study does not prove

The experiment showed improved graded performance on this assignment. It did not show improved learning.

There was no follow-up test where students had to demonstrate what they understood or retained without ChatGPT. The authors acknowledged that it remains unclear whether the GPT advantage reflected knowledge students gained or simply better output created with AI assistance.

The limits of the study also matter. It involved freshmen at one university, working on a narrow marketing task. Randomization happened across 13 class sections rather than individually across all 1,053 students.

The main performance score came from human graders who did not know which group each text belonged to. Additional measures, including causal reasoning and idea diversity, were evaluated using models from OpenAI and Anthropic.

There is also a conflict to keep in view: OpenAI provides the technology being studied and was involved in the research. Several authors work at OpenAI or were employed there during the study.

The broader lesson is not simply that students should or should not use AI. The sharper issue is whether AI supports their own thinking or replaces it. Other research cited in the source points in the same direction: homework can improve with AI while later performance without it may suffer.

For educators, the challenge is now practical. If final submissions can look stronger without proving understanding, grading has to measure the process and reasoning behind the answer, not only the finished text.