← Back to Model Beat
Research·Jul 9·all news from July 9, 2026

OpenAI finds roughly 30 percent of popular AI coding test is broken

OpenAI has identified that approximately 30 percent of the tasks in the SWE-Bench Pro coding benchmark are flawed or broken. As a result, the company has retracted its previous endorsement of the evaluation tool. This finding highlights ongoing challenges in establishing reliable metrics for measuring the programming capabilities of artificial intelligence models.

Covered by 1 source

Related stories

ResearchIncentivizing Temporal-Awareness in Egocentric Video Understanding ModelsJul 7 · 6 sourcesResearchOpenAI may have made a fatal misstep in copyright fight with news orgsJul 9 · 6 sourcesResearchLinkedIn is the undisputed king of long-form AI slop, according to a study spanning five platformsJul 12 · 4 sourcesResearchS&P Global sees OpenAI as a "key credit risk" for Oracle and cuts its credit ratingJul 12