<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj" -->

---
title: Introducing Olmo-core 3: Open, scalable training...
description: Ai2 released Olmo-core 3, a major upgrade to its open LLM training framework featuring a redesigned mixture-of-experts (MoE) training system built to scale...
canonical: https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | daily.dev
og:description: Ai2 released Olmo-core 3, a major upgrade to its open LLM training framework featuring a redesigned mixture-of-experts (MoE) training system built to scale...
og:url: https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj
og:image: https://api.daily.dev/og/posts/vzGCmhPLj.png
og:image:alt: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

**[Hugging Face](https://daily.dev/sources/huggingface)** · 7 min read · 0 upvotes · 0 comments

## Summary

Ai2 released Olmo-core 3, a major upgrade to its open LLM training framework featuring a redesigned mixture-of-experts (MoE) training system built to scale into the trillion-parameter range. It switches from FSDP to distributed data parallelism, keeping experts resident on GPUs, and adds techniques like rowwise expert parallelism, GPU-resident routing, grouped GEMM, and MXFP8 low-precision support. Benchmarks show roughly 2.7x throughput improvement over the earlier FSDP implementation, 21% gains from MXFP8 over BF16, and successful scaling tests up to 2.38 trillion total parameters. The report also documents training pitfalls such as 'token gerrymandering' in routing balance metrics. The stack underpins the next-generation Olmo model and is fully open source.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/allenai/olmocore3>

## Questions this post answers

### What throughput improvement does Olmo-core 3 offer over the previous FSDP-based MoE training implementation?

Olmo-core 3 achieves roughly 2.7x higher throughput than the earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new distributed-data-parallelism-based stack, compared to 19,400 tokens per second per GPU previously.

_Engineers scaling MoE training pipelines can follow infrastructure shifts like this one on daily.dev._

### Does using MXFP8 instead of BF16 actually speed up large model training?

Yes, enabling MXFP8 in the parts of the system where it helps most increased end-to-end training throughput by about 21% compared to BF16 in a controlled benchmark on four NVIDIA B300 GPUs, while peak active memory dropped from 103 GiB to 95 GiB. Most gains came from feed-forward computation and inter-expert data movement rather than attention.

_Teams weighing lower-precision formats for training can track results like these via daily.dev._

### What is token gerrymandering in mixture-of-experts training?

Token gerrymandering describes a failure mode where a metric designed to measure and encourage balanced expert routing improves even as the actual workload distribution across experts becomes less balanced, making the score misleading. This was identified during experiments scaling MoE training and documented as a pitfall to watch for when evaluating routing balance in large MoE systems.

_Researchers debugging MoE routing imbalances can find similar findings curated on daily.dev._

## Similar posts on daily.dev

- [Olmo 3 is a fully open LLM](https://daily.dev/posts/olmo-3-is-a-fully-open-llm-ub992gxhs) · Simon Willison · 84 upvotes · 0 comments
- [Olmo 3: Fully Open-Source LLM from AI2 \(Models, Data, & Code\)](https://daily.dev/posts/olmo-3-fully-open-source-llm-from-ai2-models-data-code--lnaxu9uge) · DigitalOcean Community · 83 upvotes · 1 comments
- [MoE inference engineering, clearly explained](https://daily.dev/posts/moe-inference-engineering-clearly-explained-pkujcfe50) · Daily Dose of Data Science \| Avi Chawla \| Substack · 0 upvotes · 0 comments
- [Accelerating Large-Scale Mixture-of-Experts Training in PyTorch](https://daily.dev/posts/accelerating-large-scale-mixture-of-experts-training-in-pytorch-yh3ady40l) · NVIDIA Developer · 0 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#gpu](https://daily.dev/tags/gpu), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs","url":"https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj"},"datePublished":"2026-10-01T15:04:17.710Z","dateModified":"2026-10-01T15:04:37.705Z","description":"Ai2 released Olmo-core 3, a major upgrade to its open LLM training framework featuring a redesigned mixture-of-experts (MoE) training system built to scale...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b844e55785860621d2593eba476f7d0d?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b844e55785860621d2593eba476f7d0d?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hugging Face","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hugging Face","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/f1f55c67d81a4330acf5b90b26b0c8e1","url":"https://daily.dev/sources/huggingface"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"machine-learning,gpu,mixture-of-experts","timeRequired":"PT7M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hugging Face","item":"https://daily.dev/sources/huggingface"},{"@type":"ListItem","position":3,"name":"Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/introducing-olmo-core-3-open-scalable-training-infrastructure-for-large-moes-vzgcmhplj#faq","mainEntity":[{"@type":"Question","name":"What throughput improvement does Olmo-core 3 offer over the previous FSDP-based MoE training implementation?","acceptedAnswer":{"@type":"Answer","text":"Olmo-core 3 achieves roughly 2.7x higher throughput than the earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new distributed-data-parallelism-based stack, compared to 19,400 tokens per second per GPU previously. Engineers scaling MoE training pipelines can follow infrastructure shifts like this one on daily.dev."}},{"@type":"Question","name":"Does using MXFP8 instead of BF16 actually speed up large model training?","acceptedAnswer":{"@type":"Answer","text":"Yes, enabling MXFP8 in the parts of the system where it helps most increased end-to-end training throughput by about 21% compared to BF16 in a controlled benchmark on four NVIDIA B300 GPUs, while peak active memory dropped from 103 GiB to 95 GiB. Most gains came from feed-forward computation and inter-expert data movement rather than attention. Teams weighing lower-precision formats for training can track results like these via daily.dev."}},{"@type":"Question","name":"What is token gerrymandering in mixture-of-experts training?","acceptedAnswer":{"@type":"Answer","text":"Token gerrymandering describes a failure mode where a metric designed to measure and encourage balanced expert routing improves even as the actual workload distribution across experts becomes less balanced, making the score misleading. This was identified during experiments scaling MoE training and documented as a pitfall to watch for when evaluating routing balance in large MoE systems. Researchers debugging MoE routing imbalances can find similar findings curated on daily.dev."}}]}
```

