<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg" -->

---
title: Ornith-1.5: An Open-Source LLM Family That Writes Its...
description: Ornith-1.5, a new open-source LLM family (9B dense, 35B MoE, 397B MoE), reportedly beats Claude Opus 4.8 on Terminal-Bench 2.1, SWE-bench Verified, WideSearch,...
canonical: https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Ornith-1.5: An Open-Source LLM Family That Writes Its Own Training Curriculum | daily.dev
og:description: Ornith-1.5, a new open-source LLM family (9B dense, 35B MoE, 397B MoE), reportedly beats Claude Opus 4.8 on Terminal-Bench 2.1, SWE-bench Verified, WideSearch,...
og:url: https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg
og:image: https://api.daily.dev/og/posts/5VrG2rEQg.png
og:image:alt: Ornith-1.5: An Open-Source LLM Family That Writes Its Own Training Curriculum
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Ornith-1.5: An Open-Source LLM Family That Writes Its Own Training Curriculum

**[Collections](https://daily.dev/sources/collections)** · 1 min read · 2 upvotes · 0 comments

## Summary

Ornith-1.5, a new open-source LLM family (9B dense, 35B MoE, 397B MoE), reportedly beats Claude Opus 4.8 on Terminal-Bench 2.1, SWE-bench Verified, WideSearch, and BrowseComp. The more notable claim is its training method: task generation, agent scaffolding, and solution generation all run inside a single RL loop using GRPO, letting the model author its own curriculum near the edge of its current ability rather than relying on human-written tasks. The piece argues that if self-generated curricula hold up, verifiable environments become more valuable than static labeled datasets for training frontier models.

## Content

Ornith-1.5 launched with a lot of buzz. It's a new open-source model family — 9B dense, 35B MoE, and a flagship 397B MoE — released under MIT license and available in vLLM from day one. The pitch: it was built through "end-to-end self-improvement," where the model generates its own training tasks, writes the scaffolding to grade them, produces solution rollouts, and feeds reward signals back through all three stages using GRPO.

That's a genuinely interesting idea. Instead of relying purely on human-curated datasets, the model searches near its own capability frontier, using a codebase or environment plus high-level instructions as a starting point. In theory, this means labs with access to rich, verifiable environments (real repositories, real tasks) could improve faster than labs just sitting on bigger static datasets. The self-generated curriculum cuts out a lot of manual task design.

And the benchmark numbers back up the hype, at least on paper. The 397B model reportedly beats Claude Opus 4.8 on four benchmarks: Terminal-Bench 2.1 (86.1 vs 85), SWE-bench Verified (86 vs 85.8), WideSearch (80.8 vs 72.9), and BrowseComp (86.6 vs 84.3). People online were calling it a

## Questions this post answers

### How does Ornith-1.5 compare to Claude Opus 4.8 on coding and search benchmarks?

The 397B MoE version of Ornith-1.5 scores 86.1 vs 85 on Terminal-Bench 2.1, 86 vs 85.8 on SWE-bench Verified, 80.8 vs 72.9 on WideSearch, and 86.6 vs 84.3 on BrowseComp compared to Claude Opus 4.8, edging ahead on all four benchmarks according to reported results.

_daily.dev surfaces benchmark comparisons like this for developers deciding which model to build agents on._

### What is a self-generated training curriculum in reinforcement learning for LLMs?

It is a training setup where the model itself proposes tasks, writes the scaffold to grade them, and generates solution rollouts, with reward propagated back through all three stages using GRPO, rather than relying on humans to author the tasks. The model searches for tasks near the edge of its current ability, removing manual curriculum design as a bottleneck.

_daily.dev helps engineers track emerging RL training techniques shaping the next wave of coding models._

---

Tags: [#ai](https://daily.dev/tags/ai), [#open-source](https://daily.dev/tags/open-source), [#llm](https://daily.dev/tags/llm), [#reinforcement-learning](https://daily.dev/tags/reinforcement-learning)

[View this post on daily.dev](https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Ornith-1.5: An Open-Source LLM Family That Writes Its Own Training Curriculum","url":"https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg"},"datePublished":"2026-08-19T14:38:11.022Z","dateModified":"2026-08-23T14:11:42.541Z","description":"Ornith-1.5, a new open-source LLM family (9B dense, 35B MoE, 397B MoE), reportedly beats Claude Opus 4.8 on Terminal-Bench 2.1, SWE-bench Verified, WideSearch,...","image":"https://pbs.twimg.com/media/HQF0-0wb0AAUOwH.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HQF0-0wb0AAUOwH.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,open-source,llm,reinforcement-learning","timeRequired":"PT1M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Ornith-1.5: An Open-Source LLM Family That Writes Its Own Training Curriculum"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/ornith-1-5-an-open-source-llm-family-that-writes-its-own-training-curriculum-5vrg2reqg#faq","mainEntity":[{"@type":"Question","name":"How does Ornith-1.5 compare to Claude Opus 4.8 on coding and search benchmarks?","acceptedAnswer":{"@type":"Answer","text":"The 397B MoE version of Ornith-1.5 scores 86.1 vs 85 on Terminal-Bench 2.1, 86 vs 85.8 on SWE-bench Verified, 80.8 vs 72.9 on WideSearch, and 86.6 vs 84.3 on BrowseComp compared to Claude Opus 4.8, edging ahead on all four benchmarks according to reported results. daily.dev surfaces benchmark comparisons like this for developers deciding which model to build agents on."}},{"@type":"Question","name":"What is a self-generated training curriculum in reinforcement learning for LLMs?","acceptedAnswer":{"@type":"Answer","text":"It is a training setup where the model itself proposes tasks, writes the scaffold to grade them, and generates solution rollouts, with reward propagated back through all three stages using GRPO, rather than relying on humans to author the tasks. The model searches for tasks near the edge of its current ability, removing manual curriculum design as a bottleneck. daily.dev helps engineers track emerging RL training techniques shaping the next wave of coding models."}}]}
```

