Aleph Alpha has released Kolibri, an open-weight bilingual model aimed at sovereign, mission-critical work in government and regulated industries. The English-German Mixture-of-Experts Transformer carries 78.1 billion parameters while activating 3.46 billion per token, supports contexts up to one million tokens, and can be downloaded with its full weights from Hugging Face under the Apache 2.0 license. Customers can run it on-premises rather than send internal data to a third-party inference service.
Small bird, fast wings, Kolibri is here.
— Aleph Alpha (@Aleph__Alpha) October 3, 2026
78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe.
Now the weights are yours. Run it on your own hardware, under Apache 2.0. pic.twitter.com/5263xZ9xZN
Kolibri builds on the earlier Kolibri Origin and Aleph Alpha's automated Model Factory. The new model was trained on 768 B200 GPUs, starting with 20 trillion tokens across 21 days, followed by mid-training and long-context adaptation for nearly 24 trillion tokens in total. Its architecture uses 384 experts with six active per token, with full attention in 10 of 50 layers and a 512-token sliding window in the other 40 to contain inference costs. Although adaptation reached 256,000 tokens, Aleph Alpha provides settings to serve the model at 1,048,576 tokens.

German accounts for 21.3% of pre-training tokens, backed by a bilingual 128,000-entry vocabulary and a tokenizer designed to preserve German compounds. The model offers four reasoning settings, none, low, medium, and high, plus tool calling. Aleph Alpha says it was specialized for German, math, coding, long-context work, and agentic tasks, with sector-specific evaluation suites for public administration, automotive, semiconductors, industrial technology, and aerospace that do not use customer data.
Dear @Aleph__Alpha team - thank you for making Kolibri-1 open.
— Konark Modi (@konarkmodi) October 3, 2026
We care deeply about sovereign AI, and launching it on German Unity Day makes today feel especially fitting.
As a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few… pic.twitter.com/mlUaFhjPCP
In the company's benchmarks, Kolibri posted an English overall score of 75.5 and a German overall score of 70.8. It scored 96.9 on AIME 2025, 85.9 on LiveCodeBench v6, and 61.4 on BFCL v4 overall. Aleph Alpha says the model sits on the quality-versus-serving-cost Pareto frontier in English and German and can match models with up to four times as many active parameters on math, code, grounding, agentic, and long-context tasks. These are vendor-run results, using Aleph Alpha's own harnesses and the highest available reasoning setting where applicable.
Grounding is central to the release. Kolibri was trained on abstention examples and Aleph Alpha's Merlin-Arthur procedure, which teaches it to withhold an answer when evidence is absent. It avoided a wrong answer on 44% of AA-Omniscience items, compared with 14.8% for Kolibri Origin, and reached 0.23 on the company's M/A grounding score. Aleph Alpha built the model in Germany, trained it in Germany and Finland, and says its control of data curation, training, evaluation, weights, and deployment is intended to meet European compliance and sovereignty requirements. Deployment uses Aleph Alpha's inference package and a Kolibri-specific vLLM plugin.
Sources and related context
- Kolibri deployment with Aleph Alpha’s vLLM plugin: Supports: The official repository documents the inference package and Kolibri-specific reasoning and tool-call parsers used to deploy the model.
- How Aleph Alpha’s Merlin-Arthur grounding method works: Related context: This research explanation describes the adversarial training procedure behind the article’s discussion of document grounding and abstention.