What a GPT-4 behavior shift means for AI monitoring

A study comparing March and June versions of GPT-3.5 and GPT-4 found major changes across several tasks, including prime-number checks and code formatting. Its authors say the results make a case for regularly monitoring language models in production, while leaving open whether the changes reflect bugs or broader quality decline.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 2 ►

The study reports less reliable behavior and less usable code formatting in newer versions, though changes varied by task and do not establish an overall quality decline.

What a GPT-4 behavior shift means for AI monitoring

A comparison of two versions of GPT-3.5 and GPT-4 found that their behavior changed substantially between March and June. The results varied by task: some capabilities appeared to improve, while others became less reliable or harder to use without extra steps.

One task showed a sharp reversal

Researchers from Stanford University and UC Berkeley evaluated the models on four kinds of work: solving math problems, answering tricky or dangerous questions, generating code, and visual thinking. The study found that the March and June versions did not behave consistently across those tasks.

The clearest example involved recognizing prime numbers. GPT-4 from March 2023 achieved 97.6% accuracy, while GPT-4 from June 2023 achieved 2.4% accuracy and ignored the chain-of-thought prompt. GPT-3.5 moved in the other direction, performing significantly better on this task in June than it had in March.

That contrast matters because a model update can affect different tasks in different ways. A result on one benchmark cannot, by itself, establish whether a model is better or worse overall.

Code became less ready to run

The study also measured whether generated code could be executed directly. For GPT-4, the share of directly executable outputs fell from 52% in March to 10% in June. For GPT-3.5, it dropped from 22% to 2%.

The researchers attributed this change to how the models followed the instruction “just the code.” In March, both models returned code that could be used directly. In June, they added triple quotes before and after the code, so users had to intervene before running it.

This finding concerns usability and output formatting. The team said the quality of the generated code appeared to be at a similar level, but it did not conduct a detailed comparison. The reported drop in direct executability therefore does not establish that the code itself had become worse.

Changes went in more than one direction

For tricky questions, GPT-4 answered fewer in June. On visual reasoning, it performed slightly better, but the June version also made errors that the March version did not. The researchers reported a slight improvement for GPT-3.5 as well.

Taken together, these results describe a mixed shift rather than a single, simple trend. The study does not give a clear answer to whether GPT-4 was worse overall in June. It does suggest that the newer version had bugs or failure patterns that were absent from the earlier one.

Why ongoing evaluation matters

The authors said the behavior of GPT-3.5 and GPT-4 varied significantly over a relatively short time. They argued that teams using language models in production should continuously evaluate how those systems behave in their own applications.

That recommendation is practical: if a model is part of a workflow, a change in its answers or formatting can affect the work around it. Monitoring can help a company notice those changes and assess whether they matter for its use case. The article says the researchers made their evaluation and ChatGPT data available on GitHub to support this kind of analysis and further research into language model drift.

The cause of the reported changes remained unsettled. The article describes possibilities including bugs, as Peter Welinder, VP Product at OpenAI, suggested in a similar example, and a general decline in quality tied to optimizations to cut costs. It says the changes were opaque to OpenAI customers, making it difficult for them to know what was behind them.

OpenAI’s Logan Kilpatrick said the company was aware of the reported regressions and was looking into them. He also called for a public OpenAI eval set that could test known regression cases as new models are released. Such evaluations could help distinguish personal reports and speculation from clearer benchmarks.

The article notes that reports of performance degradation had also surfaced around Midjourney, and frames quality control as a challenge for companies that regularly update AI models. A provider may not notice a degradation before deployment, while customers may struggle to tell whether a change has affected their work. The study’s central lesson is that model behavior needs to be measured over time, in the tasks where people actually rely on it.