← Back to Model Beat
Models·Aug 1·all news from August 1, 2026

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

Supabase has released an open source benchmarking framework called evals that tests the performance of AI coding agents against specific development tasks like schema creation and debugging. By running models such as Claude Code, Codex, and OpenCode within containerized environments, the tool provides a standardized way to measure how accurately these agents handle real-world database and infrastructure workflows. This release offers developers a transparent method to compare the reliability of AI assistants when performing complex, proprietary technical assignments.

Covered by 1 source

Related stories

ModelsDeepSeek Is Developing Massive AI Data Center in Inner MongoliaJul 29 · 61 sourcesModelsAlibaba’s Qwen3.8-Max AI Model Claims Benchmark Scores Rivaling AnthropicAug 3 · 82 sourcesModelsAnthropic AI Models Hacked Three Organizations During TestsJul 29 · 46 sourcesModelsGemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationJul 30 · 8 sources