# Kolibri 1

> Source: https://aiwiki.ai/wiki/kolibri_1
> Updated: 2026-10-06
> Fact-checked: 2026-10-06
> Categories: AI Models, Large Language Models, Mixture of Experts, Reasoning Models
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Kolibri 1." aiwiki.ai, 6 Oct 2026. https://aiwiki.ai/wiki/kolibri_1
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**Kolibri 1** is an [open-weight](https://aiwiki.ai/wiki/open_weights) [large language model](https://aiwiki.ai/wiki/large_language_model) developed by [Aleph Alpha](https://aiwiki.ai/wiki/aleph_alpha) for German and English. It supports text generation, reasoning and tool calling.[1]

Aleph Alpha released the model on 3 October 2026 for applications including document-based assistants and agentic workflows, with deployment on customer-controlled infrastructure.[2]

## Specifications

| Property | Released Kolibri 1 |
| --- | --- |
| Architecture | [Mixture-of-experts](https://aiwiki.ai/wiki/mixture_of_experts) [Transformer](https://aiwiki.ai/wiki/transformers) [3] |
| Total parameters | 78.1 billion [3] |
| Active parameters per token, excluding embeddings | 3.46 billion [3] |
| Transformer layers | 50 [3] |
| Experts per MoE layer | 384 routed experts; six selected per token plus one shared expert [3] |
| Vocabulary | 128,000 tokens [3] |
| Validated context length | 1,048,576 tokens [1] |
| Recommended context limit for serving efficiency and complex tasks | 262,144 tokens [1] |
| English and German knowledge cutoff | 18 June 2026 [1] |
| Released FP8 checkpoint | FP8 block-quantized weights; embeddings, output head, normalization layers and router remain BF16 [1] |

The model has 40 sliding-window attention layers covering the preceding 512 tokens and ten full-attention layers, with full attention in every fifth layer.[3]

Its [tokenizer](https://aiwiki.ai/wiki/tokenization), UniBPE, constructs merges bottom-up as in byte-pair encoding but selects them using a Unigram objective. Aleph Alpha designed it to represent German word structure as well as English text.[4]

## Training and development

Aleph Alpha first validated its training pipeline with **Kolibri Origin**, an internal 30.6-billion-parameter model with 3.27 billion active parameters. The release blog lists Origin as not publicly released; its specifications are separate from those of Kolibri 1.[4]

The technical report describes three base-model training stages: 20 trillion tokens of pre-training, 3.44 trillion of mid-training, and approximately 200 billion for long-context extension. German accounts for more than one fifth of the final pre-training mix.[3]

Attention and expert weights were optimized with [Muon](https://aiwiki.ai/wiki/muon_optimizer); routing used Exact Quantile Balancing to control expert load.[3]

Training used 768 NVIDIA B200 GPUs, with pre-training lasting 21 days. Post-training combined [supervised fine-tuning](https://aiwiki.ai/wiki/supervised_fine-tuning) and [reinforcement learning](https://aiwiki.ai/wiki/reinforcement_learning); the latter used over 1.2 million curated tasks covering reasoning, coding, instruction following, question answering and tools.[4]

Training sequences increased from 16,384 to 65,536 and then 262,144 tokens. The 1,048,576-token [context window](https://aiwiki.ai/wiki/context_window) was validated beyond this trained length; the recommended limit remains 262,144 tokens.[1]

## Reasoning and document grounding

Users can select reasoning effort levels `none`, `low`, `medium` and `high`, trading additional reasoning against response time and computation.[4]

Thinking can be disabled for a request using `reasoning_effort: "none"` or `enable_thinking: false`.[5]

For [retrieval-augmented generation](https://aiwiki.ai/wiki/retrieval_augmented_generation), Kolibri is trained to answer from supplied documents and abstain when the evidence is insufficient.[2]

Its training includes Aleph Alpha's Merlin-Arthur procedure.[4]

In that framework, a generator called Arthur receives modified contexts: Merlin preserves helpful evidence, while Morgana hides evidence needed to answer. The training objective rewards appropriate answers and abstention under these different conditions.[7]

The procedure's paper evaluates four other language models, from one to 32 billion parameters. Its reported improvements are evidence about the method, not a separate evaluation of Kolibri 1.[7]

The release blog reports an AA-Omniscience public-set non-hallucination rate of 44.0%, compared with 14.8% for Origin. These are benchmark-specific measurements.[4]

## Evaluation

The following are selected scores published by Aleph Alpha, rather than an independent ranking.[6]

The release blog states that evaluations used the company's harnesses and, where applicable, each model's highest reasoning effort.[4]

| Benchmark | Kolibri score |
| --- | --- |
| [AIME 2025](https://aiwiki.ai/wiki/aime_2025), English | 96.9% [6] |
| AIME 2026, English | 96.0% [6] |
| AIME 2026, German | 90.0% [6] |
| [GPQA Diamond](https://aiwiki.ai/wiki/gpqa_diamond), English | 84.3% [6] |
| GPQA Diamond, German | 81.3% [6] |
| Humanity's Last Exam, English | 21.5% [6] |
| Humanity's Last Exam, German | 15.9% [6] |
| [BFCL v4](https://aiwiki.ai/wiki/bfcl), overall | 61.4% [6] |

The technical report notes that the pre-training pool remained potentially contaminated by benchmark data and that HumanEval results were inflated by recitation.[3]

## Availability and deployment

The weights and configuration files in the official [Hugging Face](https://aiwiki.ai/wiki/hugging_face) repository are published under the [Apache License 2.0](https://aiwiki.ai/wiki/apache_license). The model card explicitly limits that grant to the artifacts included there, rather than all underlying training methods or other unpublished materials.[1]

Aleph Alpha provides the `aleph-alpha-inference` [vLLM](https://aiwiki.ai/wiki/vllm) plugin, including Kolibri's model implementation and reasoning and tool-call parsers. As of 7 October 2026, its README specifies vLLM 0.29 support. Installation also installs the supported vLLM version.[5]

The official serving example enables both parsers:[5]

```bash
pip install aleph-alpha-inference
vllm serve Aleph-Alpha/Kolibri-1 \
  --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice
```

A separate `Aleph-Alpha/Kolibri-1-BF16` checkpoint can be served without the FP8 KV-cache option.[5]

The product page lists an approximately 78 GB footprint for FP8 weights and recommends either two H100 SXM5 GPUs, two H200 GPUs, one B200 or one B300.[6]

The sparse active-parameter count does not remove the need to hold the full model in memory.[1]

## Governance and limitations

Aleph Alpha describes sovereignty in terms of control over deployment and documented model development. Its announcement reports screening training data against a blocklist of more than 4.5 million URLs and checking third-party datasets for license terms, sourcing and opt-outs.[2] These are the developer's stated governance measures, not an independent certification of every downstream application.

The model card warns of erroneous, outdated, harmful or biased outputs. It recommends application-level validation and additional safeguards for high-stakes uses, not unsupervised decision-making.[1]

## References

1. Aleph Alpha. [Kolibri 1 model card](https://huggingface.co/Aleph-Alpha/Kolibri-1). Accessed 7 October 2026.
2. Aleph Alpha. [Aleph Alpha releases Kolibri](https://aleph-alpha.com/en/news/kolibri-sovereign-ai-made-in-germany/). 5 October 2026.
3. Aleph Alpha. [Kolibri: A Sovereign European Model on the Pareto Frontier](https://aleph-alpha.com/downloads/tech-report.pdf). Technical report, especially Table 2 and Section 2.3. Accessed 7 October 2026.
4. Aleph Alpha. [Kolibri Has Landed: A Sovereign Open-Weight Model](https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/). 3 October 2026.
5. Aleph Alpha. [aleph-alpha-inference](https://github.com/Aleph-Alpha/aleph-alpha-inference). Official serving-plugin README. Accessed 7 October 2026.
6. Aleph Alpha. [Kolibri](https://aleph-alpha.com/en/kolibri/). Product specifications and evaluation tables. Accessed 7 October 2026.
7. Deiseroth, Björn, et al. [Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models](https://arxiv.org/abs/2512.11614). arXiv:2512.11614, version 3, 9 August 2026.

