Models & Research

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

· July 30, 2026
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

What changed

OpenAI says its newest API version for GPT-5.6 Sol outperforms Anthropic’s Opus 5 on the ARC-AGI-3 benchmark, but only under specific conditions. Using two extra API settings and its own testing framework, GPT-5.6 Sol scored 38.3 percent on ARC-AGI-3. However, when evaluated within the official ARC Prize environment, the same model achieved just 7.8 percent—far below Opus 5’s performance. The discrepancy raises questions about test environment fairness and API updates.

Why builders should care

Benchmark scores fluctuate depending on test setups, affecting how users pick models and vendors. OpenAI’s claim shows that upgrading API versions and tweaking settings can greatly shift model results. That matters for developers optimizing for accuracy or capabilities on complex tasks like ARC-AGI-3’s general intelligence challenges. It also signals potential instability in provider-neutral benchmarks if they don’t stay current with API versions.

The practical takeaway

Anyone relying on ARC-AGI-3 as a neutral measure should beware of outdated or inconsistent test environments that can skew model comparisons. OpenAI’s higher score depends heavily on its latest API features and non-standard settings, meaning that benchmarks alone don’t guarantee better real-world performance without full transparency. Builders and investors should demand clarity on APIs and test conditions before using benchmark claims to guide model choice or deployment.

What to watch next

Follow developments on how ARC Prize and other benchmarks update their environments and handle API versioning amid rapid model improvements. Watch whether OpenAI’s API-driven boosts become standard practice or create wider splits in reported model strength. Also, track if Anthropic or other competitors release updates that close gaps or expose new inconsistencies in these popular general AI tests.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.