Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks
AI disclosure
Summary
<p>Supabase has open sourced supabase/evals, an Apache-2.0 benchmark and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks โ building schemas, debugging Edge Functions, fixing RLS policies โ inside containerized stacks, then scores them with deterministic checks and LLM-as-a-judge.</p> <p>The post <a href="https://www.marktechpost.com/2026/08/01/supabase-releases-evals-an-open-source-benchmark-that-scores-claude-code-codex-and-opencode-on-real-supabase-tasks/">Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks</a> appeared first on <a href="https://www.marktechpost.com">MarkTechPost</a>.</p>