<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l" -->

---
title: IBM Granite 4.2: Dense Reasoning Models at 3B, 8B, and 30B
description: IBM has released Granite 4.2, a family of dense, decoder-only open-weight LLMs at 3B, 8B, and 30B parameters under Apache 2.0. Trained on ~15 trillion tokens...
canonical: https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: IBM Granite 4.2: Dense Reasoning Models at 3B, 8B, and 30B | daily.dev
og:description: IBM has released Granite 4.2, a family of dense, decoder-only open-weight LLMs at 3B, 8B, and 30B parameters under Apache 2.0. Trained on ~15 trillion tokens...
og:url: https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l
og:image: https://api.daily.dev/og/posts/0owYr0l0L.png
og:image:alt: IBM Granite 4.2: Dense Reasoning Models at 3B, 8B, and 30B
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# IBM Granite 4.2: Dense Reasoning Models at 3B, 8B, and 30B

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 2 upvotes · 0 comments

## Summary

IBM has released Granite 4.2, a family of dense, decoder-only open-weight LLMs at 3B, 8B, and 30B parameters under Apache 2.0. Trained on ~15 trillion tokens including a full trillion tokens of synthetic code, the models use a straightforward all-attention Transformer rather than the hybrid Mamba/attention designs many labs favor. The 8B and 30B variants go through GRPO-based reinforcement learning for agentic tasks, terminal use, and search, and support thinking, non-thinking, and low-effort reasoning modes plus native tool calling. Context extends to 512K tokens via a five-phase pretraining strategy, though native training is at 128K. Quantized formats (FP8, NVFP4, MXFP4, GGUF) support vLLM and llama.cpp deployment. Benchmarks are unremarkable overall, with Qwen 3.8 27B beating Granite across coding tests, but the 8B model stands out as an efficient, high-value option that runs well on consumer hardware. IBM also released two Granite Speech models the same day.

## Content

IBM has released Granite 4.2, a new family of open-weight LLMs available in 3B, 8B, and 30B parameter sizes under an Apache 2.0 license. The models are dense, decoder-only, and built on a straightforward all-attention Transformer architecture, a choice that stands out given how many labs have moved toward hybrid Mamba/attention designs. It's a continuation of the direction IBM took with Granite 4.1, which also stuck with dense models rather than following the crowd.

## Training and architecture

Each model is pre-trained from scratch on roughly 15 trillion tokens, including a full 1 trillion tokens of synthetic code generated through IBM's CodeAlchemy pipeline. Pretraining follows a five-phase strategy that pushes the context window out to 512,000 tokens, even though the models are natively trained at 128K. After pretraining, the models go through fine-tuning on chain-of-thought and agentic data.

The 8B and 30B models then go through a multi-stage reinforcement learning pipeline based on GRPO. This covers verifiable rewards, skill-specific boosters, and agentic RL targeting software engineering tasks, terminal use, and search. The result is models that support three distinct modes: thinking, non-thinking, and a low-effort mode for when you don't need the model to reason at length. Native tool calling is built in as well.

For deployment, IBM offers quantized variants in FP8, NVFP4, MXFP4, and GGUF formats, with support for vLLM and llama.cpp.

## How they actually perform

Here's where I have to be honest: the benchmarks are unremarkable. Coding performance in particular is inconsistent, and Qwen 3.8 27B beats Granite across the board on the tests that matter. That's a tough comparison to lose, especially at the 30B size.

But the 8B model deserves a mention on its own. It runs efficiently on consumer hardware and gets surprisingly close to the 30B model's performance, which makes it the more interesting release of the two for anyone who isn't running a datacenter. If you're picking between the two dense options, the 8B looks like the better value.

IBM also launched two new Granite Speech recognition models on the same day, rounding out what was clearly a coordinated release rather than a single-model drop.

Overall, Granite 4.2 doesn't move the needle on raw benchmark performance, but the reasoning modes, extended context, and agentic RL training make it a practical option for teams that want an open, dense model they can actually run and fine-tune without hybrid-architecture complications.

## Questions this post answers

### What model sizes does IBM Granite 4.2 come in and what license is it released under?

IBM Granite 4.2 is released in 3B, 8B, and 30B parameter dense models under an Apache 2.0 license. All three are decoder-only, all-attention Transformer models rather than hybrid Mamba/attention designs, continuing the dense-model direction IBM took with Granite 4.1. Deployment options include FP8, NVFP4, MXFP4, and GGUF quantized formats with vLLM and llama.cpp support.

_daily.dev surfaces open-weight model releases like this for engineers deciding what to self-host._

### How does Granite 4.2 compare to Qwen 3.8 27B on coding benchmarks?

Qwen 3.8 27B beats Granite 4.2 across the board on benchmarks that matter, and Granite's coding performance is described as inconsistent, a notable loss especially at the 30B size. Despite this, Granite 4.2's 8B model is considered the more interesting release since it runs efficiently on consumer hardware and gets close to the 30B model's performance.

_developers weighing model choices can track comparisons like this on daily.dev before committing to a stack._

### What context window does Granite 4.2 support and how was it achieved?

Granite 4.2 supports context windows up to 512,000 tokens, extended through a five-phase pretraining strategy, even though the models are natively trained at 128K tokens. Pretraining uses roughly 15 trillion tokens total, including 1 trillion tokens of synthetic code generated via IBM's CodeAlchemy pipeline, followed by fine-tuning on chain-of-thought and agentic data.

_teams evaluating long-context models for agentic workflows can follow releases like this on daily.dev._

## Similar posts on daily.dev

- [New IBM Granite 4 Models to Reduce AI Costs with Inference-Efficient Hybrid Mamba-2 Architecture](https://daily.dev/posts/new-ibm-granite-4-models-to-reduce-ai-costs-with-inference-efficient-hybrid-mamba-2-architecture-nwshaxq1n) · InfoQ · 0 upvotes · 0 comments
- [Granite 4.0 1B Speech: Compact, Multilingual, and Built for the Edge](https://daily.dev/posts/granite-4-0-1b-speech-compact-multilingual-and-built-for-the-edge-djlhogrk3) · Hugging Face · 22 upvotes · 0 comments
- [Running Granite 4 Language Models with Ollama](https://daily.dev/posts/running-granite-4-language-models-with-ollama-okn28odo9) · Niklas Heidloff · 1 upvotes · 0 comments

---

Tags: [#open-source](https://daily.dev/tags/open-source), [#llm](https://daily.dev/tags/llm), [#reinforcement-learning](https://daily.dev/tags/reinforcement-learning), [#ibm](https://daily.dev/tags/ibm)

[View this post on daily.dev](https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"IBM Granite 4.2: Dense Reasoning Models at 3B, 8B, and 30B","url":"https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l"},"datePublished":"2026-08-25T20:20:38.456Z","dateModified":"2026-09-13T19:51:25.018Z","description":"IBM has released Granite 4.2, a family of dense, decoder-only open-weight LLMs at 3B, 8B, and 30B parameters under Apache 2.0. Trained on ~15 trillion tokens...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/da00fc63851959c0e49a8b3dc2347871?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/da00fc63851959c0e49a8b3dc2347871?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"open-source,llm,reinforcement-learning,ibm","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"IBM Granite 4.2: Dense Reasoning Models at 3B, 8B, and 30B"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/ibm-granite-4-2-dense-reasoning-models-at-3b-8b-and-30b-0owyr0l0l#faq","mainEntity":[{"@type":"Question","name":"What model sizes does IBM Granite 4.2 come in and what license is it released under?","acceptedAnswer":{"@type":"Answer","text":"IBM Granite 4.2 is released in 3B, 8B, and 30B parameter dense models under an Apache 2.0 license. All three are decoder-only, all-attention Transformer models rather than hybrid Mamba/attention designs, continuing the dense-model direction IBM took with Granite 4.1. Deployment options include FP8, NVFP4, MXFP4, and GGUF quantized formats with vLLM and llama.cpp support. daily.dev surfaces open-weight model releases like this for engineers deciding what to self-host."}},{"@type":"Question","name":"How does Granite 4.2 compare to Qwen 3.8 27B on coding benchmarks?","acceptedAnswer":{"@type":"Answer","text":"Qwen 3.8 27B beats Granite 4.2 across the board on benchmarks that matter, and Granite's coding performance is described as inconsistent, a notable loss especially at the 30B size. Despite this, Granite 4.2's 8B model is considered the more interesting release since it runs efficiently on consumer hardware and gets close to the 30B model's performance. developers weighing model choices can track comparisons like this on daily.dev before committing to a stack."}},{"@type":"Question","name":"What context window does Granite 4.2 support and how was it achieved?","acceptedAnswer":{"@type":"Answer","text":"Granite 4.2 supports context windows up to 512,000 tokens, extended through a five-phase pretraining strategy, even though the models are natively trained at 128K tokens. Pretraining uses roughly 15 trillion tokens total, including 1 trillion tokens of synthetic code generated via IBM's CodeAlchemy pipeline, followed by fine-tuning on chain-of-thought and agentic data. teams evaluating long-context models for agentic workflows can follow releases like this on daily.dev."}}]}
```

