← Back to Model Beat
Models·6d ago·all news from July 30, 2026

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI reports that its GPT-5.6 Sol model achieved a 38.3 percent score on the ARC-AGI-3 benchmark, surpassing the 30.2 percent record set by Opus 5. However, this result relied on a custom test harness incorporating retained reasoning and context compaction, while the model scored only 7.8 percent in the official, standardized testing environment. This discrepancy highlights ongoing debate regarding the use of non-standard testing conditions to measure AI reasoning capabilities against established industry benchmarks.

ModelsClaude Opus 5GPT-5.6 Sol

Covered by 1 source · 2 articles

Related stories

ModelsIntroducing Claude Opus 5Jul 24 · 15 sourcesModelsKimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AIJul 16 · 247 sourcesModelsClaude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and musicAug 2ModelsGPT-5.6 Sol reportedly disproves a 30-year-old statistics conjecture in 90 minutes after humans couldn't crack itJul 15