OpenAI is challenging the current ARC-AGI-3 leaderboard narrative, but the details matter. The company says GPT-5.6 Sol can score 38.3 percent on the logic benchmark when it is run with two API settings in OpenAI's own test setup.
That number is higher than Claude Opus 5's 30.2 percent. But GPT-5.6 Sol did not reach it in the official ARC-AGI-3 harness. Under that environment, the model scored 7.8 percent.
The score depends on the harness
The central issue is not only which model gets the higher number. It is how that number is produced. OpenAI's stronger result comes from running GPT-5.6 Sol through its Responses API, rather than through the official ARC-AGI-3 environment.
OpenAI used two settings in that setup: "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it.
Those settings matter because ARC-AGI-3 tasks are action-based and depend on reasoning across steps. In the official harness, GPT-5.6 Sol's reasoning is discarded after each action, according to the source article. OpenAI links that loss of reasoning context to the much lower 7.8 percent score.
That creates two very different readings of GPT-5.6 Sol's performance. One reading says the model can reach 38.3 percent when the surrounding system preserves and compresses its reasoning history. The other says the model reaches 7.8 percent when evaluated under the official ARC-AGI-3 constraints.
Why ARC-AGI-3 makes this dispute important
ARC-AGI-3 is described as a logic benchmark. Its value comes from testing how a model performs under a defined environment, not from measuring every possible way a model can be wrapped, assisted, or orchestrated.
OpenAI argues that benchmarks never measure only the model. They also measure the technical setup around the model. That argument is important because real AI systems are usually more than a raw model call. They can include APIs, memory-like behavior, context handling, tools, and execution logic.
But the source article also points out the other side of the issue: ARC-AGI-3 deliberately tests pure model performance without external aids. That is why the official harness matters. It creates common constraints, so scores can be compared on the same basis.
When GPT-5.6 Sol is measured with OpenAI's own setup and Claude Opus 5 is measured under the official constraints, the comparison becomes harder to interpret. The 38.3 percent result may show what GPT-5.6 Sol can do in a more supportive technical environment. It does not show the same thing as the official 7.8 percent result.
How Claude Opus 5 fits into the comparison
The backdrop is Claude Opus 5's ARC-AGI-3 result. According to the source article, Anthropic's Claude Opus 5 quadrupled the record score on the benchmark and reached 30.2 percent.
That 30.2 percent result was achieved under the official constraints. The article also notes that Opus 5 would likely score even higher inside Claude Code. That point mirrors OpenAI's broader claim: the environment around a model can change outcomes.
Still, the comparison is not symmetrical. OpenAI's headline number for GPT-5.6 Sol comes from OpenAI's own harness. Claude Opus 5's cited 30.2 percent comes from the same constraints that ARC-AGI-3 uses to test pure model performance.
For readers tracking AI benchmark results, the useful distinction is simple:
- 38.3 percent is OpenAI's GPT-5.6 Sol result with its own Responses API setup.
- 7.8 percent is GPT-5.6 Sol's score in the official ARC-AGI-3 harness.
- 30.2 percent is Claude Opus 5's ARC-AGI-3 score under the official constraints.
What the numbers really say
The GPT-5.6 Sol result shows how much a benchmark score can depend on context retention and management. If a system preserves reasoning between steps and summarizes older context, the model may perform differently than it does when that reasoning is discarded after every action.
That does not make the higher number meaningless. It may be relevant for people who care about how AI systems perform when deployed with a richer execution setup. But it also does not replace the official ARC-AGI-3 score, because the official harness is designed to remove those aids.
The result is a familiar tension in AI evaluation. Developers want benchmarks that reflect real systems, while benchmark designers often want controlled tests that isolate model capability. OpenAI's argument leans toward the system view. ARC-AGI-3's official setup leans toward controlled model comparison.
For now, the most accurate takeaway is that GPT-5.6 Sol has two very different ARC-AGI-3 results depending on the harness. It beats Claude Opus 5's 30.2 percent only in OpenAI's custom setup. In the official environment, it remains far below that mark at 7.8 percent.