UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do
The United Kingdom's AI Safety Institute reports that current industry benchmarks often fail to accurately measure AI agent performance because they artificially restrict computational resources. When researchers increased token budgets during software engineering tests, agent success rates rose by 25 percent. This finding suggests that standardized testing methods currently underestimate the true technical capabilities of modern AI systems by limiting the trial-and-error processes agents use to complete complex tasks.
Covered by 1 source
- TThe Decoder↗Matthias BastianJul 3