Kolibri 1
Kolibri 1 is an open-weight large language model developed by Aleph Alpha for German and English. It supports text generation, reasoning and tool calling.[1]
Aleph Alpha released the model on 3 October 2026 for applications including document-based assistants and agentic workflows, with deployment on customer-controlled infrastructure.[2]
Specifications
| Property | Released Kolibri 1 |
|---|---|
| Architecture | Mixture-of-experts Transformer [3] |
| Total parameters | 78.1 billion [3] |
| Active parameters per token, excluding embeddings | 3.46 billion [3] |
| Transformer layers | 50 [3] |
| Experts per MoE layer | 384 routed experts; six selected per token plus one shared expert [3] |
| Vocabulary | 128,000 tokens [3] |
| Validated context length | 1,048,576 tokens [1] |
| Recommended context limit for serving efficiency and complex tasks | 262,144 tokens [1] |
| English and German knowledge cutoff | 18 June 2026 [1] |
| Released FP8 checkpoint | FP8 block-quantized weights; embeddings, output head, normalization layers and router remain BF16 [1] |
The model has 40 sliding-window attention layers covering the preceding 512 tokens and ten full-attention layers, with full attention in every fifth layer.[3]
Its tokenizer, UniBPE, constructs merges bottom-up as in byte-pair encoding but selects them using a Unigram objective. Aleph Alpha designed it to represent German word structure as well as English text.[4]
Training and development
Aleph Alpha first validated its training pipeline with Kolibri Origin, an internal 30.6-billion-parameter model with 3.27 billion active parameters. The release blog lists Origin as not publicly released; its specifications are separate from those of Kolibri 1.[4]
The technical report describes three base-model training stages: 20 trillion tokens of pre-training, 3.44 trillion of mid-training, and approximately 200 billion for long-context extension. German accounts for more than one fifth of the final pre-training mix.[3]
Attention and expert weights were optimized with Muon; routing used Exact Quantile Balancing to control expert load.[3]
Training used 768 NVIDIA B200 GPUs, with pre-training lasting 21 days. Post-training combined supervised fine-tuning and reinforcement learning; the latter used over 1.2 million curated tasks covering reasoning, coding, instruction following, question answering and tools.[4]
Training sequences increased from 16,384 to 65,536 and then 262,144 tokens. The 1,048,576-token context window was validated beyond this trained length; the recommended limit remains 262,144 tokens.[1]
Reasoning and document grounding
Users can select reasoning effort levels none, low, medium and high, trading additional reasoning against response time and computation.[4]
Thinking can be disabled for a request using reasoning_effort: "none" or enable_thinking: false.[5]
For retrieval-augmented generation, Kolibri is trained to answer from supplied documents and abstain when the evidence is insufficient.[2]
Its training includes Aleph Alpha's Merlin-Arthur procedure.[4]
In that framework, a generator called Arthur receives modified contexts: Merlin preserves helpful evidence, while Morgana hides evidence needed to answer. The training objective rewards appropriate answers and abstention under these different conditions.[7]
The procedure's paper evaluates four other language models, from one to 32 billion parameters. Its reported improvements are evidence about the method, not a separate evaluation of Kolibri 1.[7]
The release blog reports an AA-Omniscience public-set non-hallucination rate of 44.0%, compared with 14.8% for Origin. These are benchmark-specific measurements.[4]
Evaluation
The following are selected scores published by Aleph Alpha, rather than an independent ranking.[6]
The release blog states that evaluations used the company's harnesses and, where applicable, each model's highest reasoning effort.[4]
The technical report notes that the pre-training pool remained potentially contaminated by benchmark data and that HumanEval results were inflated by recitation.[3]
Availability and deployment
The weights and configuration files in the official Hugging Face repository are published under the Apache License 2.0. The model card explicitly limits that grant to the artifacts included there, rather than all underlying training methods or other unpublished materials.[1]
Aleph Alpha provides the aleph-alpha-inference vLLM plugin, including Kolibri's model implementation and reasoning and tool-call parsers. As of 7 October 2026, its README specifies vLLM 0.29 support. Installation also installs the supported vLLM version.[5]
The official serving example enables both parsers:[5]
pip install aleph-alpha-inference
vllm serve Aleph-Alpha/Kolibri-1 \
--kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
A separate Aleph-Alpha/Kolibri-1-BF16 checkpoint can be served without the FP8 KV-cache option.[5]
The product page lists an approximately 78 GB footprint for FP8 weights and recommends either two H100 SXM5 GPUs, two H200 GPUs, one B200 or one B300.[6]
The sparse active-parameter count does not remove the need to hold the full model in memory.[1]
Governance and limitations
Aleph Alpha describes sovereignty in terms of control over deployment and documented model development. Its announcement reports screening training data against a blocklist of more than 4.5 million URLs and checking third-party datasets for license terms, sourcing and opt-outs.[2] These are the developer's stated governance measures, not an independent certification of every downstream application.
The model card warns of erroneous, outdated, harmful or biased outputs. It recommends application-level validation and additional safeguards for high-stakes uses, not unsupervised decision-making.[1]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Aleph Alpha. Kolibri 1 model card. Accessed 7 October 2026.
- ^1 ^2 ^3Aleph Alpha. Aleph Alpha releases Kolibri. 5 October 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Aleph Alpha. Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, especially Table 2 and Section 2.3. Accessed 7 October 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Aleph Alpha. Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026.
- ^1 ^2 ^3 ^4Aleph Alpha. aleph-alpha-inference. Official serving-plugin README. Accessed 7 October 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Aleph Alpha. Kolibri. Product specifications and evaluation tables. Accessed 7 October 2026.
- ^1 ^2Deiseroth, Björn, et al. Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models. arXiv:2512.11614, version 3, 9 August 2026.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 1,008 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent full-article review against 7 primary and research references, October 7, 2026. Checked subject identity, specifications, availability, limitations and citation support.
Cite this page: AI Wiki. "Kolibri 1." aiwiki.ai, updated 6 Oct 2026, fact-checked 6 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/kolibri_1