Aleph Alpha's Kolibri open weights: a 78B German MoE and a vendor's own scorecard
Aleph Alpha released Kolibri, a 78B German-English open-weight MoE under Apache 2.0. We separate the vendor's self-reported scores from the verified specs.

Aleph Alpha published the full weights of Kolibri on 3 October 2026, timed to the Day of German Reunification — a deliberate date for a release the Heidelberg company wants read as European infrastructure, not a leaderboard entry. Kolibri is a bilingual German-English Mixture-of-Experts transformer: 78B total parameters, roughly 3.46B active per token, Apache 2.0. A permissive license, a full Hugging Face checkpoint, a million-token context and a bilingual tokenizer are the provable parts. The unproven part is the claim its framing leans on hardest: that it competes with models up to four times its active size.
What the release actually contains
The model repository reports 78,103,074,560 parameters in its safetensors metadata, and the config names 384 experts with 6 selected per token — matching the roughly 3.46B active per token the official post rounds down to "3B active". The context window runs to 1M tokens, though the recipe needs extra flags past 262,144: --max-model-len 1048576. The license is Apache 2.0, unusual for a vendor selling AI sovereignty, and at check time the repository showed 1,135 downloads, 385 likes and a 2 October creation date.
The scorecard is Aleph Alpha's own
The table compares Kolibri with three rival open-weight models — Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B — and every figure is the vendor's own, produced by Aleph Alpha rather than by the labs that built the rivals.
| Benchmark | Kolibri | Qwen3.6-35B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|
| AIME 2025 | 96.9 | 84.6 | 91.7 | 79.8 |
| AIME 2025 (DE) | 87.5 | 82.9 | 85.6 | 72.3 |
| AIME 2026 | 96.0 | 91.0 | 90.4 | 83.1 |
| AIME 2026 (DE) | 90.0 | 84.4 | 87.5 | 78.5 |
| GPQA diamond | 84.3 | 83.4 | 78.0 | 74.7 |
| GPQA diamond (DE) | 81.3 | 80.6 | 76.6 | 72.9 |
| AA-Omniscience | -32.8 | -15.3 | -36.5 | -24.0 |
| BrowseComp | 29.4 | 26.9 | 29.1 | — |
| tau2-bench retail | 69.9 | 71.6 | 67.5 | 62.9 |
| tau2-bench telecom | 94.7 | 99.1 | 68.1 | 41.5 |
| BFCL v4 overall | 61.4 | 67.2 | 61.0 | 58.0 |
| LiveCodeBench v6 | 85.9 | 82.5 | 82.0 | 71.2 |
| HumanEval+ | 92.7 | 92.8 | 94.7 | 92.8 |
| SWE-Bench Verified | 66.4 | — | — | — |
| LongBench Pro | 64.5 | 70.8 | 62.9 | 56.4 |
| AA-LCR | 68.3 | 69.7 | 67.0 | 52.3 |
Kolibri leads the AIME 2025 and 2026 rows and GPQA diamond in both languages and takes the two long-context rows. Its losses cluster in tool use: tau2-bench retail, tau2-bench telecom and BFCL v4 all go to Qwen3.6-35B. AA-Omniscience is a hallucination index from -100 to 100 where higher is better, and Kolibri's -32.8 is worse than Qwen3.6-35B's -15.3, so on the vendor's own honesty metric the flagship is not the best of the four. Its headline claim — matching models "up to four times its active parameter count, such as Nemotron 3 Super" — rests on AIME and GPQA, not the aggregate.
Abstention and German are the real differentiators
Two design choices explain that table. The model was trained to abstain, using a "Merlin-Arthur protocol" plus abstention data, so it answers "I don't know" when the context does not support a claim rather than confabulating. The second is language: Aleph Alpha says the custom bilingual tokenizer and a German-weighted corpus carry the (DE) rows, stating that 21.3% of pre-training tokens are German and that translation covered only about 6% of the total — worth testing, since a model leaning on machine translation differs from one trained on native text. The fuller split, roughly 62.5% English, 23.9% German and 13.6% code across 20T tokens, comes from the model card and community analysis, not the blog.
Three months separate Kolibri from its own predecessor
Kolibri was built inside what Aleph Alpha calls the Model Factory, a pipeline-as-code approach begun in January 2026. The first product, Kolibri Origin, is a 30B-total, 3B-active model with a 65k context trained on 7.5T tokens; its pre-training finished on 11 June 2026. Kolibri finished on 11 September: three months took the line from 30B to 78B parameters, 65k to 1M context and 7.5T to 20T tokens. The vendor reads those gaps as proof the pipeline works — and they are proof only that a pipeline owned end to end can be described however its owner likes.
The industrial numbers are customer proxies
Aleph Alpha also reports internal vertical results, measured by the vendor as customer-proxy suites:
| Vertical | Origin | Kolibri |
|---|---|---|
| Automotive supplier | 0.72 | 0.99 |
| Semiconductors | 0.35 | 0.80 |
| German public sector | 0.54 | 0.75 |
| Industrial drive technology | 0.31 | 0.60 |
| Aerospace | 0.14 | 0.59 |
Those jumps describe fine-tuned deployments more than a base model, and nothing about how the same tasks would score on a rival.
Compliance is framed as the product
The release is pitched at the EU AI Act, the GPAI Code of Practice and GDPR, and Aleph Alpha has signed the Code of Practice. Training ran on infrastructure in Germany and Finland, and the company says it owns the whole pipeline: data, pretraining, post-training and serving. For a European buyer whose real constraint is a regulator rather than a benchmark, that may be the load-bearing sentence.
Reception ran from respect to fax-machine jokes
On Hacker News the official-post thread reached about 650 points and 323 comments on 3 October, a second thread on a third-party explainer drew 417, and an r/LocalLLaMA thread passed 511 points the same day. "Seems like it's good at reading German documents and reasoning over them. Custom German-language tokenizer, and one of the least hallucinating models," wrote HN user pettijohn. The thread also drifted to bureaucracy jokes: HN user m00dy wrote that "A German AI model needs to be able to handle fax machines". A community claim that the run cost about $4M on a 768x B300 cluster over four weeks has no primary source. DeepSeek's open-weight launch drew the same line between vendor claims and verified specs.
What would settle it
Kolibri is a real checkpoint, a permissive license and a large bilingual model — more than most sovereign-AI announcements deliver, but the competitive claim stays unproven. Independent evaluations of the base model, third-party German-language benchmarks, tokens-per-second numbers from vLLM or llama.cpp, and a public release of the Model Factory pipeline would each move this from a press release to a measured result. Until then the honest reading is the one the license makes easy and the scorecard makes hard: the weights are open, the numbers are the vendor's own, and the case for Kolibri rests on compliance and language.


