AINews

Aleph Alpha Kolibri: the 78B open-weight model, explained

Aleph Alpha Kolibri is a 78B open-weight English and German model under Apache 2.0. Who it's for, the GPUs it needs, and what its own benchmarks prove.

Source-based. Written from the documents, reporting and reviews linked in the text. Nothing here was tested hands-on by The Ruling Desk. How we work

An Nvidia HGX B200 board with eight large GPU heatsinks and a printed label, on a display table
Photo: Pokiiri / Wikimedia Commons, CC BY-SA 4.0

Aleph Alpha Kolibri is a new open-weight model from the German AI company Aleph Alpha, released on October 3, 2026, with 78 billion parameters, two languages (English and German) and an Apache 2.0 license. The Heidelberg company pitches it for "sovereign" work in regulated fields such as public administration, industry and aerospace, where data shouldn't leave the building. Here's who it's for, what hardware it takes to run, how open the license really is, and how far Aleph Alpha's own benchmark numbers go.

Key takeaways

  • Kolibri is a mixture-of-experts model. It has 78.1 billion parameters in total, but only about 3.46 billion do the work for each token, which keeps every request cheap to compute.
  • It needs datacenter GPUs. The weights take about 78 GB, so Aleph Alpha lists one Nvidia H200, B200 or B300, or two H100 or A100 80 GB cards, as the minimum.
  • The weights are under Apache 2.0, a permissive license that allows commercial use. It covers the weights and configuration files, not Aleph Alpha's training code, methods or data.
  • Every benchmark so far is Aleph Alpha's own. Its 84.3 on GPQA Diamond tops the other mixture-of-experts models it picked by less than a point, and a dense Qwen model in the same table scores higher on most tests. We found no independent results as of October 4, 2026.

What Aleph Alpha released

Aleph Alpha published Kolibri on the Day of German Reunification, a national holiday, with a launch post on its blog and a model card on Hugging Face. The main facts, all as the company states them:

SpecKolibri
ArchitectureMixture-of-experts transformer, 50 layers, 384 experts per layer (6 routed plus 1 shared per token)
Parameters78.1B total, 3.46B active per token
LanguagesEnglish and German
Training data20T tokens of pre-training, 3.44T mid-training, 201B long-context, about 23.6T in all
Context window262,144 tokens native; validated up to 1,048,576 (1M)
ModesReasoning (none, low, medium, high) and tool calling
Knowledge cutoffJune 18, 2026
LicenseApache 2.0

A mixture-of-experts model splits its network into many small "expert" blocks, and a router sends each token to only a few of them. That's how a 78B model can run with the per-token compute of a 3B one.

The German focus is real, though the documents disagree on the exact share: the launch post says 21.3% of pre-training tokens are German, while the model card puts it at about 23.9%. Aleph Alpha trained Kolibri on 768 Nvidia B200 GPUs on infrastructure in Germany and Finland, and pre-training alone took 21 days. The technical report has the full details.

Who Aleph Alpha Kolibri is for

Aleph Alpha calls Kolibri a model "built for sovereign mission-critical work in regulated areas." In its telling, sovereignty means the model was built and trained in Europe under European and German law, with no foreign control, and that customers can run it on their own servers instead of sending internal documents to a third-party API. The company says it designed Kolibri with the EU AI Act, the GDPR and the EU's General-Purpose AI Code of Practice in mind, and the model card notes that Aleph Alpha has signed that code.

The model card is just as clear about the limits. Kolibri is meant for assistants and document workflows where a person reviews the output before anyone acts on it, not for autonomous systems, and in decision support it should advise rather than decide. Supporting two languages instead of dozens is, in the card's words, a choice of depth over breadth. If your organization works in Spanish, French or Polish, Kolibri wasn't built for you.

What hardware Kolibri needs

Mixture-of-experts cuts the compute per token, not the memory: all 78 billion parameters still have to sit in GPU memory, a trade-off the model card spells out. The published checkpoint stores its weights in FP8, an 8-bit number format, for a footprint of about 78 GB before the extra memory that long conversations need.

Aleph Alpha's own hardware guidance:

GPUMinimumRecommended
Nvidia H100 SXM522
Nvidia A100 80 GB2Not listed
Nvidia H20012
Nvidia B200 or B30011

The split follows memory. An H100 has 80 GB, barely the size of the weights, while Nvidia rates the H200 at 141 GB, enough for the weights plus room to work. Aleph Alpha also reports that two H100s served 18 concurrent 256k-token requests with Kolibri, against 3 for a 123B model it had considered. That's the company's measurement, not ours.

Two practical notes. Serving Kolibri requires Aleph Alpha's aleph-alpha-inference package, a plugin for the open-source vLLM engine, or the company's container image. And a full-precision BF16 version is also public; at 16 bits per weight, our estimate is that it needs roughly twice the memory.

How open the license really is

Apache 2.0 is one of the most permissive licenses in software: you can use, modify, fine-tune and sell products built on the model, as long as you keep the license and attribution notices (full text). No user cap or separate usage contract comes attached. The model card's "Responsible Use" section asks people to avoid unlawful uses, including practices the EU AI Act bans, but it frames that as a request and acknowledges the license is permissive.

A magnifying glass over the Hugging Face logo and address bar on a computer screen
Photo: Jernej Furman / Wikimedia Commons, CC BY 2.0

There are two limits worth knowing. The license applies only to the weights and configuration files in the Hugging Face repository, and the model card says it doesn't extend to Aleph Alpha's code, architecture, training methods or parameter settings, which the company keeps. So you get the trained model, not the means to rebuild it: Aleph Alpha describes its data curation at length, but the training data isn't part of the release.

How much the benchmarks prove

Aleph Alpha's headline scores are strong: 84.3 on GPQA Diamond (hard graduate-level science questions) in English and 81.3 in German, 96.9 on the AIME 2025 math exam and 85.9 on LiveCodeBench v6. All of them come from the company's own tests, run with its open-source eval-framework and Kolibri at its highest reasoning setting.

The full table on the model card tells a more modest story:

  • The lead is narrow. On English GPQA Diamond, Kolibri's 84.3 beats Qwen3.5 35B-A3B (83.8) and Qwen3.6 35B-A3B (83.4) by less than a point. In German, Qwen3.5 35B-A3B scores higher, 84.2 to 81.3.
  • A smaller dense model wins most rows. The same table lists Qwen3.8 27B, which scores 89.2 on GPQA Diamond and 80.2 overall in English, against Kolibri's 75.5. Aleph Alpha greys out dense models because they use several times more parameters per token, a fair point about running cost, but Qwen3.8 still scores higher.
  • 1M tokens works, with lower accuracy. On Aleph Alpha's own RULER long-context test, Kolibri's base model falls from 86.9 at 4,000 tokens to 63.2 at 1M.
  • The industry scores can't be checked. Results for the German public sector, aerospace, semiconductors and other fields come from private customer-proxy tests that outsiders can't rerun.

None of this makes the numbers wrong. They are a vendor's claims, measured on the vendor's own harness, against models the vendor chose. Because the weights are public, anyone can now run independent tests, and those are the results to watch.

Bottom line

Aleph Alpha Kolibri is a credible European open-weight option if you need English and German, want to keep data on your own hardware, and can provide one B200 or H200 class GPU or two H100s. The Apache 2.0 license is genuinely permissive for the weights, though the recipe stays with Aleph Alpha. Treat the benchmark lead as Aleph Alpha's claim until independent results arrive, and test it against Qwen's models on your own documents before you commit. It's one of several open releases this month; AWS and Cloudflare's open AI decision models take a very different approach, and the rest of our coverage lives in the AI section.

Filed under AI

Newsletter

Console and phone guides, by email.

Fixes, settings and buying decisions for the console and phone you own, from the guides we publish. Free. Unsubscribe in one click. Your email is kept by beehiiv, our newsletter service, and used only for this newsletter.