← Back to Model Beat
Models·15h ago·all news from August 1, 2026

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

Supabase has released an open source benchmarking framework called evals that tests the performance of AI coding agents against specific development tasks like schema creation and debugging. By running models such as Claude Code, Codex, and OpenCode within containerized environments, the tool provides a standardized way to measure how accurately these agents handle real-world database and infrastructure workflows. This release offers developers a transparent method to compare the reliability of AI assistants when performing complex, proprietary technical assignments.

Covered by 1 source

Related stories

ModelsAnthropic AI Models Hacked Three Organizations During TestsJul 29 · 40 sourcesModelsGemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationJul 30 · 8 sourcesModelsDeepSeek Is Developing Massive AI Data Center in Inner MongoliaJul 29 · 20 sourcesModelsAdvancing the price-performance frontier with GPT-5.6Jul 30 · 6 sources