Aleph Alpha Kolibri 1: Germany’s Sovereign 78B MoE LLM for German and English
TL;DR
Kolibri 1 is a 78 billion‑parameter mixture‑of‑experts (MoE) LLM released on 3 Oct 2026 under Apache 2.0, optimized for German and English, with a 1 million‑token context window and a design that lets it compute like a 3.5 billion‑parameter model while requiring the memory of the full 78 B model. It is built entirely on European infrastructure, complies with the EU AI Act, and is intended for on‑premise, sovereign use cases where data‑privacy and control are paramount.
What Kolibri 1 Is
| Feature | Detail |
|---|---|
| Total parameters | 78.1 B |
| Active parameters per token | 3.46 B (≈4.4 % of total) |
| Languages | German, English |
| Context length | 262 k tokens native; validated up to 1 M tokens |
| License | Apache 2.0 for weights & config (training code remains proprietary) |
| Memory footprint | ~78 GB (FP8) |
| Reasoning levels | none, low, medium, high |
| Tool calling | Supported |
| Knowledge cutoff | 18 June 2026 |
| Training data | ~24 T tokens (≈20 % German) on 768 NVIDIA B200 GPUs |
Sovereign claim – Aleph Alpha defines “sovereign” in two ways: (1) the model was built, trained, and hosted on German/Finland hardware under EU law with no foreign control, and (2) customers receive unrestricted deployment rights and IP safety, enabling on‑premise use without external lock‑in.
Core Technical Innovations
1. Mixture‑of‑Experts with 384 specialists per layer, 6 active per token
- 50 layers × 384 experts + 1 shared expert.
- Router activates 6 experts per token, reducing compute to ~3.5 B parameters while still requiring the full 78 B model to reside in memory.
"Kolibri computes like a 3.5 B parameter model but needs the memory of a 78 B model."
2. UniBPE tokenizer tuned for German compounding
- Vocabulary size: 128 k tokens.
- Uses a modified BPE scoring rule (Unigram objective) that respects German word‑formation.
- Empirical token‑count reduction: 15 % fewer tokens than GPT‑5’s
o200k_baseon the German Basic Law (35 190 vs 41 482 tokens)."Fewer tokens means fewer steps to read or write the same German text, and more German fits in the same context window."
3. Sliding‑window attention for long context
- 40 of 50 layers use 512‑token sliding‑window attention; every 5th layer uses full‑attention.
- Rotary position embeddings are only applied in sliding‑window layers, allowing context lengths beyond the 262 k training window without extra positional tricks.
- Validated up to 1 M tokens; on the RULER benchmark Kolibri scores 63.2 vs 57.5 for Qwen 3.5 35B‑A3B.
4. German‑native reasoning
- Trained on ~800 k German reasoning examples.
- German math performance: 87.5 % on AIME 2025 (German) – best among ~3 B active‑parameter models, surpassing NVIDIA Nemotron 3 Nano (84.4 %).
5. Merlin‑Arthur protocol for “I don’t know” behavior
- Three‑player game (Arthur, Merlin, Morgana) forces the model to answer only when evidence is present and to abstain otherwise.
- On the Omniscience test, Kolibri says it does not know 44 % of the time, compared with 11 % for Qwen 3.5 35B‑A3B and 23.7 % for GPT‑OSS 120B.
6. Adjustable reasoning effort
- API parameter
reasoning_effort(none,low,medium,high) lets the same model trade latency for depth on a per‑request basis.
Strengths Compared to Peer Open‑Weight Models
| Metric (≈3 B active) | Kolibri 1 | Closest competitor |
|---|---|---|
| Overall English score | 75.5 | 74.7 (Qwen 3.5 35B‑A3B) |
| Overall German score | 70.8 | 69.8 (Qwen 3.5 35B‑A3B) |
| AIME 2025 English | 96.9 | 89.6 (Nemotron 3 Nano) |
| AIME 2025 German | 87.5 | 84.4 (Nemotron 3 Nano) |
| Unseen company‑doc QA (English) | 89.7 | 87.0 (Qwen 3.5 35B‑A3B) |
| 1 M‑token context (base) | 63.2 | 58.5 (Nemotron 3 Nano) |
The model’s biggest advantage is German‑language efficiency: the tokenizer, long‑context handling, and German‑native reasoning together give it a clear edge on legal, regulatory, and technical documents written in German.
Weaknesses and Trade‑offs
- Closed‑book knowledge – ranks last among 12 evaluated models on the Retrieval‑Augmented Generation Benchmark (51.0 %); only 14.8 % correct on Omniscience questions.
- Tool‑calling – multi‑turn performance (39.8) lags behind GLM‑4.7 Flash (58.2) and Qwen 3.5 35B‑A3B (54.0).
- Coding – scores 27.7 on Terminal‑Bench 2.1, far below Qwen 3.5 35B‑A3B (39.7).
- Mid‑range context – at 128 k tokens Kolibri scores 67.9 vs Qwen 3.5’s 89.9; advantage appears only at very long windows.
- Hardware demand – ~78 GB VRAM; requires at least two 80 GB GPUs (A100/H100) or a single H200/B200/B300.
- Ecosystem maturity – needs Aleph Alpha’s custom vLLM plugin (supports vLLM 0.29 only at launch); no hosted inference providers yet.
- Language scope – limited to German and English by design.
- Dense‑model competition – Qwen 3.8 27B (dense) outperforms Kolibri on standard English/German benchmarks while using ~8× more active parameters per token.
Running Kolibri 1
# Install Aleph Alpha’s inference plugin (includes the required vLLM version)
pip install 'aleph-alpha-inference>=1'
# Serve with FP8 KV cache; enable tool calling and reasoning parser
vllm serve Aleph-Alpha/Kolibri-1 \
--kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
The server presents an OpenAI‑compatible endpoint. Example Python client:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="Aleph-Alpha/Kolibri-1",
messages=[{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts‑Modell ist."}],
extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "enable_thinking": True}},
temperature=1.0, top_p=0.97, top_k=128,
)
print(resp.choices[0].message.content)
For the 1 M‑token context, add:
vllm serve Aleph-Alpha/Kolibri-1 \
--kv-cache-dtype fp8 \
--max-model-len 1048576 \
--hf-overrides '{"max_position_embeddings": 1048576}'
Ideal Use Cases
- German‑centric enterprises – banks, automotive suppliers, aerospace firms, or public authorities that must keep data on‑premise and require high‑quality German output.
- Long‑document RAG – legal statutes, contracts, manuals, or medical guidelines where a 1 M‑token context and token‑efficient German tokenizer reduce latency and cost.
- Safety‑critical applications – scenarios where “I don’t know” is preferable to hallucination, such as clinical decision support or regulatory compliance checks.
- Sovereign AI strategy – organizations seeking EU‑law‑compliant models with full deployment freedom and IP protection.
When Not to Choose Kolibri
- Projects that need strong closed‑book trivia or coding assistance.
- Multilingual workloads beyond German/English.
- Environments without access to ≥78 GB GPU memory.
- Applications that rely heavily on multi‑turn tool‑calling or need the best mid‑range context performance.
Community Reaction Highlights
"I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models." – niemandhier
"The absence of any comparison to Qwen3.8 Flash, another MoE model with a small‑ish (6B) number of active parameters, is pretty striking." – spijdar
"Nice to see public goods in this space." – veryfancy
"A bigger dense model beats it. Qwen3.8 27B ... How is a 27B dense model bigger than a 78B MoE?" – woadwarrior01
These comments underline two themes: the strategic value of a sovereign, auditable model for regulated sectors, and the need for broader benchmark comparisons (especially against newer dense models) to fully assess Kolibri’s trade‑offs.
Bottom Line
Kolibri 1 demonstrates that a European‑built, open‑weight MoE can deliver state‑of‑the‑art German language performance while keeping compute costs low through expert routing. Its strengths lie in long‑context German RAG, honest abstention, and sovereign deployment. The model’s high memory requirement, limited language support, and weaker closed‑book knowledge mean it is a niche solution for organizations that prioritize data sovereignty and German‑centric workloads over raw breadth or coding prowess.
Sources
Related
- Dispatch
- Project
- Dispatch
- Dispatch
- Dispatch