← Back to Model Beat
Research·Jul 3·all news from July 3, 2026

UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do

The United Kingdom's AI Safety Institute reports that current industry benchmarks often fail to accurately measure AI agent performance because they artificially restrict computational resources. When researchers increased token budgets during software engineering tests, agent success rates rose by 25 percent. This finding suggests that standardized testing methods currently underestimate the true technical capabilities of modern AI systems by limiting the trial-and-error processes agents use to complete complex tasks.

Covered by 1 source

Related stories

ResearchIncentivizing Temporal-Awareness in Egocentric Video Understanding ModelsJul 7 · 6 sourcesResearchTaming Text-to-Sounding Video Generation via Advanced Modality Condition and InteractionJul 7ResearchDSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot ManipulationJul 7 · 15 sourcesResearchWeblica: Scalable and Reproducible Training Environments for Visual Web AgentsJul 7