OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness
OpenAI reports that its GPT-5.6 Sol model achieved a 38.3 percent score on the ARC-AGI-3 benchmark, surpassing the 30.2 percent record set by Opus 5. However, this result relied on a custom test harness incorporating retained reasoning and context compaction, while the model scored only 7.8 percent in the official, standardized testing environment. This discrepancy highlights ongoing debate regarding the use of non-standard testing conditions to measure AI reasoning capabilities against established industry benchmarks.
Covered by 1 source · 2 articles
- TThe Decoder↗Matthias Bastian6d ago
- TThe Decoder↗Matthias Bastian6d ago