← Back to Model Beat
Models·Jul 30·all news from July 30, 2026

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI reports that its GPT-5.6 Sol model achieved a 38.3 percent score on the ARC-AGI-3 benchmark, surpassing the 30.2 percent record set by Opus 5. However, this result relied on a custom test harness incorporating retained reasoning and context compaction, while the model scored only 7.8 percent in the official, standardized testing environment. This discrepancy highlights ongoing debate regarding the use of non-standard testing conditions to measure AI reasoning capabilities against established industry benchmarks.

ModelsClaude Opus 5GPT-5.6 Sol

Covered by 1 source · 2 articles

Related stories

ModelsIntroducing Claude Opus 5Jul 24 · 15 sourcesModelsKimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AIJul 16 · 252 sourcesModelsPreviewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speedAug 13 · 5 sourcesModelsImproving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free usersAug 6 · 5 sources