OpenAI finds roughly 30 percent of popular AI coding test is broken
OpenAI has identified that approximately 30 percent of the tasks in the SWE-Bench Pro coding benchmark are flawed or broken. As a result, the company has retracted its previous endorsement of the evaluation tool. This finding highlights ongoing challenges in establishing reliable metrics for measuring the programming capabilities of artificial intelligence models.
Covered by 1 source
- TThe Decoder↗Maximilian SchreinerJul 9