Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
Artificial Analysis updated its Intelligence Index to version 4.2 following criticism regarding how it evaluated the performance of GPT-6 Astra. The revised scoring places GPT-6 Astra four points ahead of its previous iteration, though the model remains behind Claude Fable 5.1. This adjustment addresses concerns over the accuracy of the platform's benchmarking methodology for evaluating modern AI capabilities.
ModelsGPT-6 Astra
Covered by 3 sources · 7 articles
- TThe Decoder↗Matthias BastianSep 4
- TThe Decoder↗Matthias BastianSep 5
- TThe Decoder↗Maximilian SchreinerSep 4
- HHacker News↗wertykSep 3
- HHacker News↗cebertSep 5
- GGizmodo↗Sep 4
- HHacker News↗zof3Sep 5