<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf" -->

---
title: Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash...
description: Claude Fable 5.1 matched Claude Opus 5&#x27;s accuracy on the Agents on Rails leaderboard, solving 92% of runs, while beating it on price, speed, and security-task...
canonical: https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha) | daily.dev
og:description: Claude Fable 5.1 matched Claude Opus 5&#x27;s accuracy on the Agents on Rails leaderboard, solving 92% of runs, while beating it on price, speed, and security-task...
og:url: https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf
og:image: https://api.daily.dev/og/posts/ENrlM6NrF.png
og:image:alt: Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)

**[Rails](https://daily.dev/sources/rails)** · 3 min read · 0 upvotes · 0 comments

## Summary

Claude Fable 5.1 matched Claude Opus 5's accuracy on the Agents on Rails leaderboard, solving 92% of runs, while beating it on price, speed, and security-task handling. It also jumped Rails API recall from prior scores up to 41%. The previously stealth 'ox-alpha' model has been revealed as GLM 5.3 Flash from Z.ai, scoring 83% at roughly one-fifteenth the cost of Grok 4.6, which matched the same score. Benchmarks used the lemans harness with default effort levels across 21 atomic Writebook tasks, three attempts each, with results published on the Agents on Rails page.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://rubyonrails.org/2026/9/2/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash>

## Questions this post answers

### How does Claude Fable 5.1 compare to Claude Opus 5 on Rails coding tasks?

Claude Fable 5.1 matched Claude Opus 5's accuracy, solving 58 of 63 runs (92%) and passing 20 of 21 tasks at least once, tying for the top spot on the Agents on Rails leaderboard. It beat Opus 5 on every other metric, running about 40% cheaper, faster, and handling a security pen-test task correctly where the prior Fable 5 failed all three attempts.

_Comparing coding agent releases before switching models is easier with benchmark coverage on daily.dev._

### What model was hidden behind the ox-alpha stealth name in Rails AI benchmarks?

The stealth model called ox-alpha turned out to be GLM 5.3 Flash from Z.ai. Rerunning all 63 attempts under its real name matched the pre-release results, scoring 52 of 63 (83%) at a total cost of $3.31, roughly five cents per run, matching Grok 4.6's 83% score at about one-fifteenth the cost.

_Developers tracking which cheap models rival pricier ones can follow model-pricing shifts on daily.dev._

### What is Rails API recall and how did Claude Fable 5.1 score on it?

Rails API recall measures whether a coding agent's solution uses the specific idiomatic Rails API expected, rather than a hand-rolled workaround that still passes tests. Prior models scored between 8% and 35%, but Claude Fable 5.1 reached 41%, notably recognizing long-standing APIs like quote_column_name that no earlier model had mentioned unprompted.

_Engineers judging whether an AI agent writes idiomatic Rails code can track these benchmarks on daily.dev._

## Similar posts on daily.dev

- [Agents on Rails: the first benchmark report](https://daily.dev/posts/agents-on-rails-the-first-benchmark-report-1bsnzp6hz) · RUBYLAND · 2 upvotes · 2 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#ai-agents](https://daily.dev/tags/ai-agents), [#rails](https://daily.dev/tags/rails), [#claude](https://daily.dev/tags/claude), [#glm](https://daily.dev/tags/glm)

[View this post on daily.dev](https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)","url":"https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf"},"datePublished":"2026-09-02T12:14:23.692Z","dateModified":"2026-09-02T12:38:35.002Z","description":"Claude Fable 5.1 matched Claude Opus 5's accuracy on the Agents on Rails leaderboard, solving 92% of runs, while beating it on price, speed, and security-task...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f3219107ad4885c4e776650b7cdd8928?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f3219107ad4885c4e776650b7cdd8928?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Rails","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Rails","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/d44501fa43a2441a9e27a08367cdfa52","url":"https://daily.dev/sources/rails"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,ai-agents,rails,claude,glm","timeRequired":"PT3M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Rails","item":"https://daily.dev/sources/rails"},{"@type":"ListItem","position":3,"name":"Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha--enrlm6nrf#faq","mainEntity":[{"@type":"Question","name":"How does Claude Fable 5.1 compare to Claude Opus 5 on Rails coding tasks?","acceptedAnswer":{"@type":"Answer","text":"Claude Fable 5.1 matched Claude Opus 5's accuracy, solving 58 of 63 runs (92%) and passing 20 of 21 tasks at least once, tying for the top spot on the Agents on Rails leaderboard. It beat Opus 5 on every other metric, running about 40% cheaper, faster, and handling a security pen-test task correctly where the prior Fable 5 failed all three attempts. Comparing coding agent releases before switching models is easier with benchmark coverage on daily.dev."}},{"@type":"Question","name":"What model was hidden behind the ox-alpha stealth name in Rails AI benchmarks?","acceptedAnswer":{"@type":"Answer","text":"The stealth model called ox-alpha turned out to be GLM 5.3 Flash from Z.ai. Rerunning all 63 attempts under its real name matched the pre-release results, scoring 52 of 63 (83%) at a total cost of $3.31, roughly five cents per run, matching Grok 4.6's 83% score at about one-fifteenth the cost. Developers tracking which cheap models rival pricier ones can follow model-pricing shifts on daily.dev."}},{"@type":"Question","name":"What is Rails API recall and how did Claude Fable 5.1 score on it?","acceptedAnswer":{"@type":"Answer","text":"Rails API recall measures whether a coding agent's solution uses the specific idiomatic Rails API expected, rather than a hand-rolled workaround that still passes tests. Prior models scored between 8% and 35%, but Claude Fable 5.1 reached 41%, notably recognizing long-standing APIs like quote_column_name that no earlier model had mentioned unprompted. Engineers judging whether an AI agent writes idiomatic Rails code can track these benchmarks on daily.dev."}}]}
```

