Aleph Alpha Kolibri: the 78B open-weight model, explained
Aleph Alpha Kolibri is a 78B open-weight English and German model under Apache 2.0. Who it's for, the GPUs it needs, and what its own benchmarks prove.
Source-based. Written from the documents, reporting and reviews linked in the text. Nothing here was tested hands-on by The Ruling Desk. How we work

Aleph Alpha Kolibri is a new open-weight model from the German AI company Aleph Alpha, released on October 3, 2026, with 78 billion parameters, two languages (English and German) and an Apache 2.0 license. The Heidelberg company pitches it for "sovereign" work in regulated fields such as public administration, industry and aerospace, where data shouldn't leave the building. Here's who it's for, what hardware it takes to run, how open the license really is, and how far Aleph Alpha's own benchmark numbers go.
Key takeaways
- Kolibri is a mixture-of-experts model. It has 78.1 billion parameters in total, but only about 3.46 billion do the work for each token, which keeps every request cheap to compute.
- It needs datacenter GPUs. The weights take about 78 GB, so Aleph Alpha lists one Nvidia H200, B200 or B300, or two H100 or A100 80 GB cards, as the minimum.
- The weights are under Apache 2.0, a permissive license that allows commercial use. It covers the weights and configuration files, not Aleph Alpha's training code, methods or data.
- Every benchmark so far is Aleph Alpha's own. Its 84.3 on GPQA Diamond tops the other mixture-of-experts models it picked by less than a point, and a dense Qwen model in the same table scores higher on most tests. We found no independent results as of October 4, 2026.
What Aleph Alpha released
Aleph Alpha published Kolibri on the Day of German Reunification, a national holiday, with a launch post on its blog and a model card on Hugging Face. The main facts, all as the company states them:
| Spec | Kolibri |
|---|---|
| Architecture | Mixture-of-experts transformer, 50 layers, 384 experts per layer (6 routed plus 1 shared per token) |
| Parameters | 78.1B total, 3.46B active per token |
| Languages | English and German |
| Training data | 20T tokens of pre-training, 3.44T mid-training, 201B long-context, about 23.6T in all |
| Context window | 262,144 tokens native; validated up to 1,048,576 (1M) |
| Modes | Reasoning (none, low, medium, high) and tool calling |
| Knowledge cutoff | June 18, 2026 |
| License | Apache 2.0 |
A mixture-of-experts model splits its network into many small "expert" blocks, and a router sends each token to only a few of them. That's how a 78B model can run with the per-token compute of a 3B one.
The German focus is real, though the documents disagree on the exact share: the launch post says 21.3% of pre-training tokens are German, while the model card puts it at about 23.9%. Aleph Alpha trained Kolibri on 768 Nvidia B200 GPUs on infrastructure in Germany and Finland, and pre-training alone took 21 days. The technical report has the full details.
Who Aleph Alpha Kolibri is for
Aleph Alpha calls Kolibri a model "built for sovereign mission-critical work in regulated areas." In its telling, sovereignty means the model was built and trained in Europe under European and German law, with no foreign control, and that customers can run it on their own servers instead of sending internal documents to a third-party API. The company says it designed Kolibri with the EU AI Act, the GDPR and the EU's General-Purpose AI Code of Practice in mind, and the model card notes that Aleph Alpha has signed that code.
The model card is just as clear about the limits. Kolibri is meant for assistants and document workflows where a person reviews the output before anyone acts on it, not for autonomous systems, and in decision support it should advise rather than decide. Supporting two languages instead of dozens is, in the card's words, a choice of depth over breadth. If your organization works in Spanish, French or Polish, Kolibri wasn't built for you.
What hardware Kolibri needs
Mixture-of-experts cuts the compute per token, not the memory: all 78 billion parameters still have to sit in GPU memory, a trade-off the model card spells out. The published checkpoint stores its weights in FP8, an 8-bit number format, for a footprint of about 78 GB before the extra memory that long conversations need.
Aleph Alpha's own hardware guidance:
| GPU | Minimum | Recommended |
|---|---|---|
| Nvidia H100 SXM5 | 2 | 2 |
| Nvidia A100 80 GB | 2 | Not listed |
| Nvidia H200 | 1 | 2 |
| Nvidia B200 or B300 | 1 | 1 |
The split follows memory. An H100 has 80 GB, barely the size of the weights, while Nvidia rates the H200 at 141 GB, enough for the weights plus room to work. Aleph Alpha also reports that two H100s served 18 concurrent 256k-token requests with Kolibri, against 3 for a 123B model it had considered. That's the company's measurement, not ours.
Two practical notes. Serving Kolibri requires Aleph Alpha's aleph-alpha-inference package, a plugin for the open-source vLLM engine, or the company's container image. And a full-precision BF16 version is also public; at 16 bits per weight, our estimate is that it needs roughly twice the memory.
How open the license really is
Apache 2.0 is one of the most permissive licenses in software: you can use, modify, fine-tune and sell products built on the model, as long as you keep the license and attribution notices (full text). No user cap or separate usage contract comes attached. The model card's "Responsible Use" section asks people to avoid unlawful uses, including practices the EU AI Act bans, but it frames that as a request and acknowledges the license is permissive.

There are two limits worth knowing. The license applies only to the weights and configuration files in the Hugging Face repository, and the model card says it doesn't extend to Aleph Alpha's code, architecture, training methods or parameter settings, which the company keeps. So you get the trained model, not the means to rebuild it: Aleph Alpha describes its data curation at length, but the training data isn't part of the release.
How much the benchmarks prove
Aleph Alpha's headline scores are strong: 84.3 on GPQA Diamond (hard graduate-level science questions) in English and 81.3 in German, 96.9 on the AIME 2025 math exam and 85.9 on LiveCodeBench v6. All of them come from the company's own tests, run with its open-source eval-framework and Kolibri at its highest reasoning setting.
The full table on the model card tells a more modest story:
- The lead is narrow. On English GPQA Diamond, Kolibri's 84.3 beats Qwen3.5 35B-A3B (83.8) and Qwen3.6 35B-A3B (83.4) by less than a point. In German, Qwen3.5 35B-A3B scores higher, 84.2 to 81.3.
- A smaller dense model wins most rows. The same table lists Qwen3.8 27B, which scores 89.2 on GPQA Diamond and 80.2 overall in English, against Kolibri's 75.5. Aleph Alpha greys out dense models because they use several times more parameters per token, a fair point about running cost, but Qwen3.8 still scores higher.
- 1M tokens works, with lower accuracy. On Aleph Alpha's own RULER long-context test, Kolibri's base model falls from 86.9 at 4,000 tokens to 63.2 at 1M.
- The industry scores can't be checked. Results for the German public sector, aerospace, semiconductors and other fields come from private customer-proxy tests that outsiders can't rerun.
None of this makes the numbers wrong. They are a vendor's claims, measured on the vendor's own harness, against models the vendor chose. Because the weights are public, anyone can now run independent tests, and those are the results to watch.
Bottom line
Aleph Alpha Kolibri is a credible European open-weight option if you need English and German, want to keep data on your own hardware, and can provide one B200 or H200 class GPU or two H100s. The Apache 2.0 license is genuinely permissive for the weights, though the recipe stays with Aleph Alpha. Treat the benchmark lead as Aleph Alpha's claim until independent results arrive, and test it against Qwen's models on your own documents before you commit. It's one of several open releases this month; AWS and Cloudflare's open AI decision models take a very different approach, and the rest of our coverage lives in the AI section.