Aleph Alpha releases open-weight Kolibri with 1M context

Kolibri is now on Hugging Face with 78.1B total parameters, 3.46B active per token, tool calling and support for contexts up to one million tokens.

· 3 min read
Aleph Alpha

Aleph Alpha has released Kolibri, an open-weight bilingual model aimed at sovereign, mission-critical work in government and regulated industries. The English-German Mixture-of-Experts Transformer carries 78.1 billion parameters while activating 3.46 billion per token, supports contexts up to one million tokens, and can be downloaded with its full weights from Hugging Face under the Apache 2.0 license. Customers can run it on-premises rather than send internal data to a third-party inference service.

Kolibri builds on the earlier Kolibri Origin and Aleph Alpha's automated Model Factory. The new model was trained on 768 B200 GPUs, starting with 20 trillion tokens across 21 days, followed by mid-training and long-context adaptation for nearly 24 trillion tokens in total. Its architecture uses 384 experts with six active per token, with full attention in 10 of 50 layers and a 512-token sliding window in the other 40 to contain inference costs. Although adaptation reached 256,000 tokens, Aleph Alpha provides settings to serve the model at 1,048,576 tokens.

Aleph Alpha

German accounts for 21.3% of pre-training tokens, backed by a bilingual 128,000-entry vocabulary and a tokenizer designed to preserve German compounds. The model offers four reasoning settings, none, low, medium, and high, plus tool calling. Aleph Alpha says it was specialized for German, math, coding, long-context work, and agentic tasks, with sector-specific evaluation suites for public administration, automotive, semiconductors, industrial technology, and aerospace that do not use customer data.

In the company's benchmarks, Kolibri posted an English overall score of 75.5 and a German overall score of 70.8. It scored 96.9 on AIME 2025, 85.9 on LiveCodeBench v6, and 61.4 on BFCL v4 overall. Aleph Alpha says the model sits on the quality-versus-serving-cost Pareto frontier in English and German and can match models with up to four times as many active parameters on math, code, grounding, agentic, and long-context tasks. These are vendor-run results, using Aleph Alpha's own harnesses and the highest available reasoning setting where applicable.

Grounding is central to the release. Kolibri was trained on abstention examples and Aleph Alpha's Merlin-Arthur procedure, which teaches it to withhold an answer when evidence is absent. It avoided a wrong answer on 44% of AA-Omniscience items, compared with 14.8% for Kolibri Origin, and reached 0.23 on the company's M/A grounding score. Aleph Alpha built the model in Germany, trained it in Germany and Finland, and says its control of data curation, training, evaluation, weights, and deployment is intended to meet European compliance and sovereignty requirements. Deployment uses Aleph Alpha's inference package and a Kolibri-specific vLLM plugin.

Source