AI benchmarks have a trust problem and Google wants to fix it
Google DeepMind has launched a pilot program using double-blind testing to evaluate a frontier AI model. By leveraging Confidential Space technology, the process prevents researchers from accessing model weights while ensuring Google cannot see the specific test prompts. This approach aims to address widespread concerns regarding the reliability of AI benchmarks, which are currently susceptible to data contamination and bias. Establishing a secure, neutral framework could provide a more credible standard for measuring the capabilities and safety of future artificial intelligence systems.
Covered by 1 source
- TThe Decoder↗Maximilian SchreinerAug 28