← Back to Model Beat
Research·Aug 28·all news from August 28, 2026

AI benchmarks have a trust problem and Google wants to fix it

Google DeepMind has launched a pilot program using double-blind testing to evaluate a frontier AI model. By leveraging Confidential Space technology, the process prevents researchers from accessing model weights while ensuring Google cannot see the specific test prompts. This approach aims to address widespread concerns regarding the reliability of AI benchmarks, which are currently susceptible to data contamination and bias. Establishing a secure, neutral framework could provide a more credible standard for measuring the capabilities and safety of future artificial intelligence systems.

Covered by 1 source

Related stories

ResearchUS Department of Justice backs fair use for AI training in landmark copyright caseSep 1 · 7 sourcesResearchMetaRoCE: A New RDMA Transport Built for AI-Scale EthernetAug 24 · 2 sourcesResearchFederal Appeals Court Blocks Charge Over Private Possession of AI-Generated Child Sexual Abuse ImagesAug 30 · 5 sourcesResearchFrontier Red Team ResearchSep 1 · 2 sources