An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Researchers introduced the MirrorCode benchmark to test an AI model's ability to reconstruct full software applications without access to the original source code. While top-performing models like Claude Opus successfully rebuilt large codebases, current AI still struggles to complete the most complex programming challenges.
Covered by 1 source
- TThe Decoder↗Matthias BastianJun 26