Models & Research

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

· July 30, 2026
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

What happened

OpenAI claims its GPT-5.6 Sol model outperforms Anthropic’s Opus 5 on the ARC-AGI-3 benchmark, scoring 38.3 percent compared to Opus 5’s 30.2 percent. However, this result only holds when testing GPT-5.6 Sol through OpenAI’s own custom API setup that retains reasoning and compresses context. In the official ARC-AGI-3 test environment, GPT-5.6 Sol managed just 7.8 percent accuracy, far lower than Opus 5’s reported performance without any special aids.

Why it matters

This split result reveals the importance of benchmarking context and testing conditions. OpenAI’s score boost depends on a proprietary test harness that keeps intermediate reasoning steps and compacts input context, enhancing performance artificially compared to the standardized ARC-AGI-3 environment. For AI buyers, investors, and builders, this means not all benchmark claims translate into genuine, out-of-the-box model capability. The industry is pressured to demand transparent, uniform testing standards to prevent inflated performance figures skewing competitive positioning.

What to watch next

Watch for how OpenAI and Anthropic respond in refining benchmark protocols or disclosing testing methods. Investors and customers should track if GPT-5.6 Sol’s advantage holds up in real-world applications or remains tied to bespoke testing frameworks. The wider field will monitor whether benchmarks evolve to account for these kinds of context retention tricks or whether similar tactics become widespread, complicating honest comparison between large AI models.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.