← Back to Model Beat
Models·Jun 26·all news from June 26, 2026

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run

Researchers introduced the MirrorCode benchmark to test an AI model's ability to reconstruct full software applications without access to the original source code. While top-performing models like Claude Opus successfully rebuilt large codebases, current AI still struggles to complete the most complex programming challenges.

Covered by 1 source

Related stories

ModelsOpenAI Leans Toward Waiting Until 2027 for IPO, NYT SaysJun 25 · 25 sourcesModelsClaude Science, an AI workbench for scientists, is now availableJun 30 · 12 sourcesModelsAnthropic’s Mythos 5 AI Model Cleared by US for Wider UseJun 26 · 11 sourcesModelsDeepseek's DSpark boosts AI speed by up to 85 percent, a strategic win under tightening US export controlsJun 27 · 12 sources