<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/introducing-supabase-evals-sydt4bsio" -->

---
title: Introducing Supabase Evals | daily.dev
description: Supabase has open-sourced supabase/evals, a benchmark and framework for measuring how well AI coding agents (Claude Code, Codex, OpenCode) perform on real...
canonical: https://daily.dev/posts/introducing-supabase-evals-sydt4bsio
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Introducing Supabase Evals | daily.dev
og:description: Supabase has open-sourced supabase/evals, a benchmark and framework for measuring how well AI coding agents (Claude Code, Codex, OpenCode) perform on real...
og:url: https://daily.dev/posts/introducing-supabase-evals-sydt4bsio
og:image: https://api.daily.dev/og/posts/SYdT4bsIO.png
og:image:alt: Introducing Supabase Evals
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Introducing Supabase Evals

**[Supabase](https://daily.dev/sources/supabase)** · 6 min read · 14 upvotes · 0 comments

## Summary

Supabase has open-sourced supabase/evals, a benchmark and framework for measuring how well AI coding agents (Claude Code, Codex, OpenCode) perform on real Supabase tasks like schema building, Edge Function debugging, and RLS policy fixes. The framework runs agents against actual Supabase environments using containerized stacks, scoring results with deterministic checks and LLM-as-a-judge. Key findings include: agents perform reasonably well without skills loaded, but skills help with edge cases and outdated knowledge; agents tend to avoid declarative schema workflows; newer Supabase libraries like @supabase/server are underused; skill activation is uneven across models; and Claude Code checks docs far less frequently than Codex-based agents. The benchmark results are publicly viewable and the regression suite runs daily internally.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://supabase.com/blog/introducing-supabase-evals>

## Questions this post answers

### What is supabase/evals and what does it measure?

Supabase/evals is an open-source benchmark and framework that tests how well AI coding agents such as Claude Code, Codex, and OpenCode build with Supabase. It runs scenarios like building a schema, debugging a failed Edge Function, or fixing a broken RLS policy against real hosted-like and local CLI Supabase environments, scoring results with deterministic checks and LLM-as-a-judge.

_Developers weighing which AI agent to trust with database work can follow benchmarks like this one on daily.dev._

### Why do AI coding agents avoid using declarative schema workflows in Supabase projects?

Agents tend to hand-write migration files even in projects already using declarative schemas, which describe a database's shape in one place instead of across multiple migration files. This happens even when declarative schemas are the preferred workflow, prompting Supabase to update its skill guidance to clarify when each approach should be used and to verify the fix through evals.

_Teams debugging inconsistent agent behavior around schema management can track fixes like this on daily.dev._

### Why aren't AI agents using the new @supabase/server package for writing Edge Functions?

Agents tasked with writing Edge Functions still verify auth manually with supabase-js instead of adopting the newer @supabase/server package meant to simplify secure Edge Function boilerplate, because the new library isn't sufficiently discoverable. Supabase responded by publishing a 'Which package to choose' guide to help agents and developers pick the right package.

_Developers deciding between overlapping libraries can find guidance like this through daily.dev._

## Similar posts on daily.dev

- [AI Agents Know About Supabase. They Don't Always Use It Right.](https://daily.dev/posts/ai-agents-know-about-supabase-they-don-t-always-use-it-right--jdyn6pgoc) · Supabase · 41 upvotes · 3 comments
- [AWS Releases Aws-Bench to Evaluate Agents on Cloud Tasks](https://daily.dev/posts/aws-releases-aws-bench-to-evaluate-agents-on-cloud-tasks-0shw5bgj3) · InfoQ · 1 upvotes · 0 comments
- [Introducing o11y-bench: an open benchmark for AI agents running observability workflows](https://daily.dev/posts/introducing-o11y-bench-an-open-benchmark-for-ai-agents-running-observability-workflows-psucsa908) · Grafana Labs · 20 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#mcp](https://daily.dev/tags/mcp), [#supabase](https://daily.dev/tags/supabase)

[View this post on daily.dev](https://daily.dev/posts/introducing-supabase-evals-sydt4bsio)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Introducing Supabase Evals","url":"https://daily.dev/posts/introducing-supabase-evals-sydt4bsio","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/introducing-supabase-evals-sydt4bsio"},"datePublished":"2026-07-31T19:12:56.601Z","dateModified":"2026-09-13T18:41:55.689Z","description":"Supabase has open-sourced supabase/evals, a benchmark and framework for measuring how well AI coding agents (Claude Code, Codex, OpenCode) perform on real...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/e425d46ae9e79eb3bb6168384585ae57?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/e425d46ae9e79eb3bb6168384585ae57?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Supabase","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Supabase","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/d0aa77cfb88d44dd8e6d5cd02fdf80bd","url":"https://daily.dev/sources/supabase"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/introducing-supabase-evals-sydt4bsio","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":14},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,mcp,supabase","timeRequired":"PT6M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Supabase","item":"https://daily.dev/sources/supabase"},{"@type":"ListItem","position":3,"name":"Introducing Supabase Evals"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/introducing-supabase-evals-sydt4bsio#faq","mainEntity":[{"@type":"Question","name":"What is supabase/evals and what does it measure?","acceptedAnswer":{"@type":"Answer","text":"Supabase/evals is an open-source benchmark and framework that tests how well AI coding agents such as Claude Code, Codex, and OpenCode build with Supabase. It runs scenarios like building a schema, debugging a failed Edge Function, or fixing a broken RLS policy against real hosted-like and local CLI Supabase environments, scoring results with deterministic checks and LLM-as-a-judge. Developers weighing which AI agent to trust with database work can follow benchmarks like this one on daily.dev."}},{"@type":"Question","name":"Why do AI coding agents avoid using declarative schema workflows in Supabase projects?","acceptedAnswer":{"@type":"Answer","text":"Agents tend to hand-write migration files even in projects already using declarative schemas, which describe a database's shape in one place instead of across multiple migration files. This happens even when declarative schemas are the preferred workflow, prompting Supabase to update its skill guidance to clarify when each approach should be used and to verify the fix through evals. Teams debugging inconsistent agent behavior around schema management can track fixes like this on daily.dev."}},{"@type":"Question","name":"Why aren't AI agents using the new @supabase/server package for writing Edge Functions?","acceptedAnswer":{"@type":"Answer","text":"Agents tasked with writing Edge Functions still verify auth manually with supabase-js instead of adopting the newer @supabase/server package meant to simplify secure Edge Function boilerplate, because the new library isn't sufficiently discoverable. Supabase responded by publishing a 'Which package to choose' guide to help agents and developers pick the right package. Developers deciding between overlapping libraries can find guidance like this through daily.dev."}}]}
```

