How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
OpenAI researchers significantly increased the performance of GPT-5.6 on the ARC-AGI-3 benchmark by adjusting two specific API settings. These modifications improved the model’s reasoning capabilities and data compaction, demonstrating that architectural fine-tuning can yield substantial efficiency gains without requiring fundamental changes to the underlying model.
Covered by 1 source
- OOpenAI Blog↗15h ago