<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc" -->

---
title: IBM Granite 4.1: how the 8B model beats a 32B MoE
description: IBM has released Granite 4.1, its largest model family to date, featuring dense decoder-only language models at 3B, 8B, and 30B parameters trained on ~15...
canonical: https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: IBM Granite 4.1: how the 8B model beats a 32B MoE | daily.dev
og:description: IBM has released Granite 4.1, its largest model family to date, featuring dense decoder-only language models at 3B, 8B, and 30B parameters trained on ~15...
og:url: https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc
og:image: https://api.daily.dev/og/posts/6DKXUjOKC.png
og:image:alt: IBM Granite 4.1: how the 8B model beats a 32B MoE
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# IBM Granite 4.1: how the 8B model beats a 32B MoE

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 41 upvotes · 1 comments

## Summary

IBM has released Granite 4.1, its largest model family to date, featuring dense decoder-only language models at 3B, 8B, and 30B parameters trained on ~15 trillion tokens under Apache 2.0. The standout result is the 8B instruct model matching or outperforming the previous 32B MoE model across benchmarks like ArenaHard, BFCL V3, and GSM8K. Training used a five-phase pre-training pipeline, 4.1M fine-tuning samples filtered via LLM-as-Judge, and a four-stage RLHF pipeline using on-policy GRPO with DAPO loss. Long-context support reaches 512K tokens on the 8B and 30B models via staged context extension with model merging. The family also includes Granite Vision 4.1, Granite Speech 4.1 (5.33% WER), Granite Guardian 4.1 for safety moderation, and a multilingual embedding model supporting 200+ languages. All models are available on Hugging Face and IBM watsonx, with support for vLLM, SGLang, llama.cpp, and Ollama.

## Content

IBM has released Granite 4.1, its largest model family to date. The lineup includes dense decoder-only language models at 3B, 8B, and 30B parameter sizes, all trained on roughly 15 trillion tokens and released under Apache 2.0.

The headline result is the 8B instruct model matching or outperforming the previous-generation Granite 4.0-H-Small, a 32B mixture-of-experts model, across nearly every benchmark including ArenaHard, BFCL V3, and GSM8K. That's a meaningful efficiency jump.

## How they built it

Training used a five-phase pre-training pipeline with evolving data mixtures at each stage. Fine-tuning drew from roughly 4.1 million samples curated through an LLM-as-Judge filtering system, which IBM used to screen quality before any RL work began.

The reinforcement learning pipeline runs four stages using on-policy GRPO with DAPO loss. One detail worth noting: mid-training, IBM caught a math performance regression caused by RLHF and corrected it before it compounded. That kind of mid-run diagnosis is harder than it sounds.

Long-context support reaches 512K tokens on the 8B and 30B models via staged context extension combined with model merging, which preserves short-context performance rather than trading it away.

## The rest of the family

Granite 4.1 isn't just the language models. The release also includes:

- **Granite Vision 4.1** for document understanding, covering tables, charts, and key-value pair extraction
- **Granite Speech 4.1**, hitting 5.33% word error rate, with a non-autoregressive variant designed for higher throughput
- **Granite Guardian 4.1** for safety moderation and harm detection
- **Granite Embedding Multilingual R2** supporting 200+ languages

## Where to get them

All models are available on Hugging Face and IBM's watsonx platform. The language models run on vLLM, SGLang, llama.cpp, and Ollama. FP8 quantized variants are available for memory-constrained deployments.

## Community discussion

Top comments from developers on daily.dev.

**@yaireo** · 1 upvotes

> Very interesting: [https://research.ibm.com/blog/granite-4-1-ai-foundation-models](https://research.ibm.com/blog/granite-4-1-ai-foundation-models)
>
> I wonder if it will ever be available on _Cursor_

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#llm](https://daily.dev/tags/llm), [#reinforcement-learning](https://daily.dev/tags/reinforcement-learning), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"IBM Granite 4.1: how the 8B model beats a 32B MoE","url":"https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc"},"datePublished":"2026-05-03T09:16:47.307Z","dateModified":"2026-05-03T09:17:23.743Z","description":"IBM has released Granite 4.1, its largest model family to date, featuring dense decoder-only language models at 3B, 8B, and 30B parameters trained on ~15...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/ee8bcab9438804e5c35aff87146d9200?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/ee8bcab9438804e5c35aff87146d9200?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":41},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"machine-learning,llm,reinforcement-learning,mixture-of-experts","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"IBM Granite 4.1: how the 8B model beats a 32B MoE"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/ibm-granite-4-1-how-the-8b-model-beats-a-32b-moe-6dkxujokc","comment":[{"@type":"Comment","text":"Very interesting: https://research.ibm.com/blog/granite-4-1-ai-foundation-models\nI wonder if it will ever be available on Cursor","datePublished":"2026-05-04T13:50:56.381Z","url":"https://daily.dev/posts/6DKXUjOKC#c-AuoWm9QyT","author":{"@type":"Person","name":"Yair Even Or","url":"https://daily.dev/yaireo","image":"https://avatars.githubusercontent.com/u/845031?v=4"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}}]}
```

