<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv" -->

---
title: Aleph Alpha releases Kolibri, an open-weight German...
description: Aleph Alpha released Kolibri, an open-weight European language model with 78B total parameters but only about 3.46B active per token via a mixture-of-experts...
canonical: https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Aleph Alpha releases Kolibri, an open-weight German model with 78B parameters and 3.5B active | daily.dev
og:description: Aleph Alpha released Kolibri, an open-weight European language model with 78B total parameters but only about 3.46B active per token via a mixture-of-experts...
og:url: https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv
og:image: https://api.daily.dev/og/posts/BVFAfp6Tv.png
og:image:alt: Aleph Alpha releases Kolibri, an open-weight German model with 78B parameters and 3.5B active
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Aleph Alpha releases Kolibri, an open-weight German model with 78B parameters and 3.5B active

**[Collections](https://daily.dev/sources/collections)** · 3 min read · 0 upvotes · 0 comments

## Summary

Aleph Alpha released Kolibri, an open-weight European language model with 78B total parameters but only about 3.46B active per token via a mixture-of-experts design (384 experts per layer, 6 active). It supports up to 1M tokens of context using a mostly-local attention pattern with periodic global-attention layers. It tops open models of its size in English and German benchmarks, scores 96.9% on AIME, and uses a German-optimized tokenizer needing 15% fewer tokens than GPT-5's for German text. It was trained with roughly 800k German reasoning examples and a hide-the-evidence game that improved its ability to say 'I don't know' (44% vs Qwen3.5's 11%). Weaknesses include memory-based QA, long multi-turn tool use, and coding agents, where Qwen models lead. Running it requires two H100s or one H200 and Aleph Alpha's vLLM add-on.

## Content

Aleph Alpha, a German AI lab, has released Kolibri, an open-weight model under the Apache 2.0 license. It has 78B parameters, of which 3.46B are active per token, and a context window of up to 1M tokens. The team announced it on German National Day. Aside from Mistral, it may be the first general-purpose model from an EU lab to reach this level of performance.

The model supports an explicit reasoning mode and tool calling, and it is optimized for long-context and inference efficiency.

## How it works

**Mixture of experts.** Each layer has 384 small experts, and a router sends every token to 6 of them. Compute per token is about that of a 3.5B model, but the full 78B has to sit in memory. That is roughly 78 GB, so you need two H100s or one H200.

**Long context on a budget.** Most layers look only at the last 512 tokens, and every fifth layer looks at everything. That is how it reaches 1M tokens, a few thick books' worth, without the cost exploding.

**A German tokenizer.** Aleph Alpha built a bilingual German/English tokenizer, and 21.3% of the pre-training tokens are organic German data. @TejasKumar_ ran the tokenizer on the German constitution and found it needed 15% fewer tokens than GPT-5's. "Bundesverfassungsgericht" is 6 tokens for GPT-5 and 2 for Kolibri. Fewer tokens means cheaper, faster German, and more text fits in the context window.

**Reasoning in German.** According to @TejasKumar_, the team found that a small amount of German reasoning data is worse than none, because the model's German chains of thought go in circles and never finish. So they built about 800k German reasoning examples and trained on a lot of them.

**Saying "I don't know".** Training includes a game where parts of the documents are hidden, sometimes to help the model and sometimes to hide the evidence, and the model has to tell which. When it didn't know an answer, Kolibri admitted it 44% of the time. Qwen3.5 did so 11% of the time.

## Results

In Aleph Alpha's evals, Kolibri tops every open model of its size in English and German. It scores 96.9% on AIME, ahead of every mixture-of-experts model tested, including ones three times larger. Only a dense model doing 8x the work beats it.

## Weak spots

Qwen models are ahead on answering from memory, on using tools over long multi-turn exchanges, and on coding agents. Running Kolibri also requires Aleph Alpha's add-on for vLLM, the open-source model server.

## Compliance

Aleph Alpha says Kolibri was built with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the start. Copyright law was a particular focus.

## Who it's for

If you have German documents that must stay on your own hardware, this is a strong option. Anyone can run it on their own servers. It's another sign that more EU labs are entering the open-model race.

## Questions this post answers

### What are the specs of Aleph Alpha's new Kolibri model and how does it compare to other open models?

Kolibri is an open-weight mixture-of-experts model with 78B total parameters but only about 3.46B active per token, supporting up to 1M tokens of context. It uses 384 experts per layer with 6 routed per token, mostly local attention (last 512 tokens) with every fifth layer attending globally, and requires two H100s or one H200 to run. It scores 96.9% on AIME, beating larger MoE models, and tops open models of its size in English and German.

_Developers evaluating self-hosted German-language models can track releases like this on daily.dev._

### How does Kolibri's tokenizer compare to GPT-5's for German text?

Kolibri's tokenizer needs about 15% fewer tokens than GPT-5's tokenizer for the same German text, such as the German constitution. For example, the word 'Bundesverfassungsgericht' takes 6 tokens with GPT-5's tokenizer but only 2 with Kolibri's, making German text processing cheaper, faster, and allowing more content to fit in the context window.

_Teams comparing tokenizer efficiency across models can follow benchmarks like this on daily.dev._

### What are the main weaknesses of Aleph Alpha's Kolibri model compared to Qwen models?

Kolibri is weaker than Qwen models at answering from memory, at tool use across long multi-turn exchanges, and at coding agent tasks. It also requires Aleph Alpha's own add-on for vLLM to run, rather than working out of the box with the standard open-source model server.

_Developers choosing between open-weight models for agentic coding can weigh trade-offs like these on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 3 discussions and 32 comments across x (as of 2026-10-03).

**TL;DR:** Replies are largely enthusiastic about Kolibri's efficiency and German-language focus, with practical clarifications on hardware requirements and MoE tradeoffs dominating the substantive discussion.

**Sentiment:** 55% positive · 40% mixed · 5% skeptical

**The case for**

- The model's 'I don't know' behavior on unsupported contexts is seen as a standout and meaningful strength.
- People are excited to deploy it into real apps and appreciate having a strong model originating from Europe.
- Some note the AIME math score is impressive for the model's size.

**The pushback**

- Commenters point out that despite only 3.5B active parameters, all 78B must still fit in memory, so it still requires data-center-grade GPUs, not a laptop.
- One person questions whether a high benchmark score like AIME actually translates to solving real customer tasks.
- Someone questions the company's claimed European/German sovereignty given it's largely Canadian-owned.

**By community**

- x (positive): Replies skew enthusiastic, mixing hype with substantive clarifications about memory requirements, quantization, and the model's calibrated uncertainty behavior.

**Open questions**

- How does the model perform in practice for real customer-facing tasks beyond high benchmark scores like AIME?
- How does it actually compare against current frontier closed models?

**Highlights**

> @TejasKumar_ @Aleph__Alpha worth splitting the two numbers: 3.5b per word is the compute bill, 78b total is the memory bill. sparsity buys speed, not a smaller deployment. size the box for 78 and enjoy the price of 3.5.
> — [johnroodepic on x · 1 points, 1 comments](https://x.com/johnroodepic/status/2106326680537313388)

> @AIQuanting yes, 78 gb is fp8, so bf16 would be about 156. and one 80 gb h100 doesn't quite cut it even at fp8, since the kv cache and activations need room on top of the weights. that's why the model card's minimum is 2 h100s or 1 h200
> — [TejasKumar\_ on x · 1 comments](https://x.com/TejasKumar_/status/2106365759752421449)

> @TejasKumar_ @Aleph__Alpha Great breakdown. The number I'm proudest of isn't in the math column though: Kolibri says "I don't know" when the context doesn't support an answer.
> — [IlhanScheer on x · 3 points, 2 comments](https://x.com/IlhanScheer/status/2106317979134640426)

> @TejasKumar_ @Aleph__Alpha 96.9% on AIME is a model card. I hire for the seat that owns the call when that score still fails the customer task.
> — [genedai on x](https://x.com/genedai/status/2106314259504517629)

> @TejasKumar_ @Aleph__Alpha Well I'm all for European/German sovereignty but isn't this company 90% Canadian
> — [goldenagesurfer on x](https://x.com/goldenagesurfer/status/2106423914679210211)

**Source threads**

- [x](https://x.com/ClementDelangue/status/2106348150684303422) · 0 points · 0 comments
- [x](https://x.com/TejasKumar_/status/2106311074270363939) · 0 points · 0 comments
- [x](https://x.com/TejasKumar_/status/2106310341793583167) · 0 points · 32 comments

## Similar posts on daily.dev

- [German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German](https://daily.dev/posts/german-ai-consortium-releases-soofi-s-an-open-30b-model-that-tops-benchmarks-in-both-english-and-ge-hrv00052t) · Hacker News · 0 upvotes · 0 comments
- [Z.ai pitches GLM-5.2 for long-running software engineering tasks](https://daily.dev/posts/z-ai-pitches-glm-5-2-for-long-running-software-engineering-tasks-3x3e7fy5z) · InfoWorld · 1 upvotes · 0 comments

---

Tags: [#open-source](https://daily.dev/tags/open-source), [#vllm](https://daily.dev/tags/vllm), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Aleph Alpha releases Kolibri, an open-weight German model with 78B parameters and 3.5B active","url":"https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv"},"datePublished":"2026-10-03T11:38:46.396Z","dateModified":"2026-10-03T18:36:22.793Z","description":"Aleph Alpha released Kolibri, an open-weight European language model with 78B total parameters but only about 3.46B active per token via a mixture-of-experts...","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"open-source,vllm,mixture-of-experts","timeRequired":"PT3M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Aleph Alpha releases Kolibri, an open-weight German model with 78B parameters and 3.5B active"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/aleph-alpha-releases-kolibri-an-open-weight-german-model-with-78b-parameters-and-3-5b-active-bvfafp6tv#faq","mainEntity":[{"@type":"Question","name":"What are the specs of Aleph Alpha's new Kolibri model and how does it compare to other open models?","acceptedAnswer":{"@type":"Answer","text":"Kolibri is an open-weight mixture-of-experts model with 78B total parameters but only about 3.46B active per token, supporting up to 1M tokens of context. It uses 384 experts per layer with 6 routed per token, mostly local attention (last 512 tokens) with every fifth layer attending globally, and requires two H100s or one H200 to run. It scores 96.9% on AIME, beating larger MoE models, and tops open models of its size in English and German. Developers evaluating self-hosted German-language models can track releases like this on daily.dev."}},{"@type":"Question","name":"How does Kolibri's tokenizer compare to GPT-5's for German text?","acceptedAnswer":{"@type":"Answer","text":"Kolibri's tokenizer needs about 15% fewer tokens than GPT-5's tokenizer for the same German text, such as the German constitution. For example, the word 'Bundesverfassungsgericht' takes 6 tokens with GPT-5's tokenizer but only 2 with Kolibri's, making German text processing cheaper, faster, and allowing more content to fit in the context window. Teams comparing tokenizer efficiency across models can follow benchmarks like this on daily.dev."}},{"@type":"Question","name":"What are the main weaknesses of Aleph Alpha's Kolibri model compared to Qwen models?","acceptedAnswer":{"@type":"Answer","text":"Kolibri is weaker than Qwen models at answering from memory, at tool use across long multi-turn exchanges, and at coding agent tasks. It also requires Aleph Alpha's own add-on for vLLM to run, rather than working out of the box with the standard open-source model server. Developers choosing between open-weight models for agentic coding can weigh trade-offs like these on daily.dev."}}]}
```

