Kolibri-1 is an open-weight language model that Aleph Alpha released on October 3, 2026. It has 78.1 billion parameters, activates 3.46 billion of them per token, and is built for German and English. It supports adjustable reasoning and tool calling, and its weights are available on Hugging Face under the Apache 2.0 license. Aleph Alpha trained the model on infrastructure in Germany and Finland. This guide covers the model's specifications, license, benchmark results, hardware requirements, and community quantizations, and shows how to run a self-hosted Kolibri-1 endpoint behind production AI agents.
Kolibri-1 at a glance
Kolibri-1 is a mixture-of-experts model: each layer holds many expert networks, and a router sends each token to a small subset of them. Compute per token stays close to that of a 3.5B model, while the full network stores 78.1B parameters. The specifications below come from the official Kolibri-1 model card and Aleph Alpha's release announcement.
| Specification | Kolibri-1 |
|---|---|
| Developer | Aleph Alpha, Heidelberg, Germany |
| Release date | October 3, 2026 |
| Parameters | 78.1B total, 3.46B active per token |
| Architecture | Mixture of experts, 50 layers, 384 routed experts plus 1 shared expert per layer, 6 routed experts per token |
| Attention | Sliding-window attention over 512 tokens in four of every five layers, full attention in the fifth |
| Context window | 262,144 tokens native, validated up to 1,048,576 |
| Languages | German and English |
| Reasoning | Adjustable effort: none, low, medium, high |
| Tool calling | Yes, through the official vLLM parser |
| Weights | FP8 (about 78 GB) and BF16 (about 156 GB) |
| Knowledge cutoff | June 18, 2026 |
| Pre-training data | 20 trillion tokens, 21.3% of them German (about 4.3 trillion), per the release announcement |
| Pre-training compute | 768 NVIDIA B200 GPUs for 21 days |
| Training location | Infrastructure in Germany and Finland |
| License | Apache 2.0 for the published weights and configuration files |
Aleph Alpha also trained a new tokenizer that follows German word structure, including compound nouns. The model card reports 4.7 bytes per token on German text, a higher compression than leading models with larger vocabularies, while English stays at 4.2 bytes per token. Fewer tokens for the same German text mean lower latency and serving cost per request.
The Apache 2.0 license allows commercial use, modification, and redistribution of the weights. Its scope is defined narrowly: according to the model card, the rights apply to the weights and configuration files in the Hugging Face repository and do not extend to code, model architecture, parameter settings, or training methods. The serving plugin, aleph-alpha-inference, is published separately under Apache 2.0.
Kolibri-1 benchmarks
Aleph Alpha compared the post-trained model with 13 other models, mostly using its open-source eval-framework with the same setup for every model and Kolibri-1 at reasoning effort high. The comparison groups mixture-of-experts models with similar active parameter counts and lists dense models, which activate several times more parameters per token, for reference. All figures below are Aleph Alpha's own measurements, not independent results.
Compared with similar open models
The table shows Kolibri-1 next to three mixture-of-experts models of similar active size and Qwen3.8 27B, the higher-scoring of the two dense models in Aleph Alpha's comparison.
| Benchmark | Kolibri-1 | Qwen3.6 35B-A3B | Gemma 4 26B-A4B | Mistral Small 4 119B-A6B | Qwen3.8 27B (dense) |
|---|---|---|---|---|---|
| Overall, English | 75.5 | 71.4 | 71.9 | 63.1 | 80.2 |
| Overall, German | 70.8 | 67.3 | 66.3 | 61.4 | 79.9 |
| Agentic average (EN) | 63.4 | 62.1 | 54.6 | 40.7 | 66.7 |
| Tau3-Bench Banking | 38.1 | 10.6 | 16.0 | 5.7 | 50.0 |
| BFCL v4 multi-turn | 47.5 | 58.1 | 61.4 | 40.4 | 55.5 |
| Agentic Wiki QA (DE) | 69.4 | 58.8 | 66.2 | 65.4 | 73.5 |
| Industry RAG average (DE) | 67.5 | 65.8 | 49.4 | 53.4 | 80.2 |
| LiveCodeBench v6 | 85.9 | 82.5 | 82.3 | 71.2 | 93.8 |
| AA-Omniscience non-hallucination rate | 44.0 | 56.7 | 14.3 | 34.7 | 67.3 |
Selected post-training results from the Kolibri-1 model card (FP8 release), October 2026. Kolibri-1 ran at reasoning effort high; the other models used their documented settings. Bold marks the best mixture-of-experts score in each row; the shaded column is the dense reference model.
Kolibri-1 posts the highest overall English and German scores of every mixture-of-experts model in the comparison, but the margins are narrow. Qwen3.5 35B-A3B reaches 74.7 in English and 69.8 in German, and GPT-OSS 120B scores 70.2 in German. Qwen3.8 27B leads Kolibri-1 by about 5 points in English and 9 in German, while activating almost eight times as many parameters per token.
Agent and tool-calling results
Kolibri-1's English agentic average of 63.4 ties Qwen3.5 35B-A3B for the best mixture-of-experts result. Its clearest lead is on Tau3-Bench Banking, a simulated banking customer-service task, with 38.1 against 16.0 or less for every other mixture-of-experts model. It reaches 76.7 on Tau2-Bench Airline, the best mixture-of-experts score there, and 69.4 on Aleph Alpha's German agentic Wiki QA task, also the best in its group.
Multi-turn function calling is the weak spot. On the multi-turn split of the Berkeley Function Calling Leaderboard (BFCL v4), Kolibri-1 scores 47.5, below 9 of the 11 other mixture-of-experts models, and its overall BFCL v4 score of 61.4 trails Qwen3.6 35B-A3B (67.2) and Gemma 4 26B-A4B (68.2). On TerminalBench 2.1, a command-line agent benchmark, it scores 27.7 against 39.7 for Qwen3.5 35B-A3B. Agents that chain many dependent tool calls within one conversation should be tested on exactly that pattern before rollout.
Grounding and hallucinations
Aleph Alpha trained Kolibri-1 with abstention data so that it declines questions it cannot answer reliably. On the RGB negative-rejection test, which checks whether a model declines when the supplied documents contain no answer, it scores 85.6. Its AA-Omniscience non-hallucination rate of 44.0 trails Qwen3.6 35B-A3B (56.7) but is far ahead of Gemma 4 26B-A4B (14.3). Correcting false information in supplied documents is weaker: on the RGB fact-check test it scores 34.0, while the mixture-of-experts models from other developers range from 58.0 to 90.0. Retrieval agents built on Kolibri-1 therefore benefit from checks that compare answers against their sources.
Context window: 1M validated, 262k recommended
Kolibri-1's final long-context training phase used 262,144-token sequences, which is its native context length. Positional encoding sits only in the sliding-window layers, so the context can grow further without position scaling, and Aleph Alpha validated quality and serving efficiency up to 1,048,576 tokens. The model card still recommends at most 262,144 tokens for latency- or throughput-sensitive deployments and for complex tasks. The default configuration serves 262,144 tokens; longer contexts need the additional vLLM flags --max-model-len 1048576 and --hf-overrides '{"max_position_embeddings": 1048576}'.
Long-context results for the base model are consistent with that recommendation. Its RULER average falls from 86.9 at 4k tokens to 69.8 at 256k and 63.2 at 1M. At 1M tokens it scores higher than the base versions of Nemotron 3 Nano 30B-A3B (58.5) and Qwen3.5 35B-A3B (57.5), but both of them score higher at every length up to 512k.
Agents rarely need a million tokens in a single request. A long-running agent accumulates tool results and conversation history, and keeping each request under the recommended window protects both latency and answer quality. Connic's context compression shortens the history once prompts reach a configured token limit; the guide to context compression explains the trade-offs.
Hardware requirements and serving
The sparse design reduces compute per token, but all 78.1B parameters must stay in GPU memory. Aleph Alpha lists the following configurations; the KV cache and runtime overhead come on top of the weights.
| Weights | Weight memory | Minimum | Recommended |
|---|---|---|---|
| FP8 (Kolibri-1) | About 78 GB | 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300 | 2× H100 SXM5, 2× H200, 1× B200 or 1× B300 |
| BF16 (Kolibri-1-BF16) | About 156 GB | 4× A100 80 GB, 4× H100 SXM5, 2× H200, 1× B200 or 1× B300 | 4× H100 SXM5, 2× H200, 2× B200 or 1× B300 |
Serving requires Aleph Alpha's aleph-alpha-inference package, which adds Kolibri-1 support to vLLM and installs the vLLM version it supports. A prebuilt container image is available as well. The official command enables reasoning and tool calling:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choiceAleph Alpha recommends temperature 1.0, top_p 0.97, and top_k 128, and the repository ships these values in its generation config. Reasoning effort is set per request through the chat template, with the levels none, low, medium, and high. When a request sets no effort, the template uses high, which adds reasoning tokens and latency to every call. Tool calls use a Hermes-style format that the kolibri1 parser converts into structured tool calls in the OpenAI-compatible response, and tool calling works together with reasoning.
Quantized versions: GGUF, MLX, and NVFP4
Aleph Alpha publishes only the FP8 and BF16 checkpoints. Within a day of the release, community members uploaded more than a dozen GGUF, MLX, and NVFP4 conversions to Hugging Face. As of October 4, 2026, the GGUF and MLX files need patched or forked runtimes. A selection:
| Format | Repository | Size | Runtime | Notes |
|---|---|---|---|---|
| GGUF Q3_K_S | Eliasfpv28/Kolibri-1-Q3_K_S-GGUF | 33.9 GB | Patched llama.cpp | Started locally at a 4,096-token context; reasoning and tool calling not validated |
| GGUF Q4_K_M, Q8_0 | Hob-forge/Kolibri-1-GGUF | 47.5 GB, 83.1 GB | Patched llama.cpp | Architecture patch checked against a reference implementation |
| MLX 3-bit | here-be-dragons-ai/Kolibri-1-MLX-3bit | 33 GiB | mlx-vlm fork | Sized for Macs with 48 GB unified memory |
| MLX 2-bit | velaia/Kolibri-1-MLX-2bit | 24 GiB | mlx-lm with a bundled launcher | Macs with 36 GB or more |
| MLX 4-bit | velaia/Kolibri-1-MLX-4bit | 41 GiB | mlx-lm with a bundled launcher | Macs with 64 GB or more |
| MLX 6-bit | audreyt/Kolibri-1-MLX-6bit | 59.16 GiB | mlx-lm fork | Macs with 96 GB or more |
| NVFP4 W4A16 | audreyt/Kolibri-1-NVFP4-W4A16 | 44.14 GiB | vLLM | Full generation not yet validated |
Stock llama.cpp, Ollama, and LM Studio cannot load the GGUF files until they support the kolibri1 architecture. On October 4, a llama.cpp feature request and an mlx-lm pull request for Kolibri-1 were still open. The NVFP4 build targets vLLM with Aleph Alpha's plugin but has not completed a generation test yet. Most conversions start from the FP8 checkpoint, and their authors report smoke tests and small perplexity samples rather than benchmark results.
What Kolibri-1 means for Europe
A German company developed Kolibri-1, and according to the release announcement it was trained in Germany and Finland under European and German law. The Apache 2.0 weights run in any EU data center or on premises, so prompts and documents stay on infrastructure the operator controls. Teams comparing hosting options can read our overview of EU-hosted AI models.
The release documentation also addresses EU rules for general-purpose AI models. The model card states that Aleph Alpha is a signatory of the EU General-Purpose AI Code of Practice and names a contact for rightsholders; its responsible-use guidance asks users not to apply the model to practices prohibited under Article 5 of the EU AI Act. It describes the intended deployment as assistants and agentic workflows in which a person reviews the output before it is acted on, rather than systems that act unreviewed. Obligations for a deployed agent still depend on its use case; the guide to EU AI Act compliance for AI agents covers that assessment.
German and English are the only supported languages. That makes Kolibri-1 a candidate for German-language workloads in Germany, Austria, and Switzerland and for European companies serving German-speaking customers, but not for products that need the EU's other languages. Aleph Alpha positions the model for public administration, aerospace, manufacturing, automotive, and semiconductor companies.
How to use Kolibri-1 with Connic
Connic runs AI agents in production with deployments, connectors, guardrails, approvals, judges, and traces built in. Kolibri-1 is not part of the Connic-managed model catalog, but a self-hosted Kolibri-1 endpoint connects as a custom OpenAI-compatible provider and then works with all of these features like any other model. The setup takes five steps.
1. Serve the model
Run the official vLLM command from the previous section on a GPU server that meets the hardware requirements. If the endpoint is reachable from outside its network, start vLLM with --api-key so that only clients with the key can call it.
2. Add a custom provider
In Connic, open Project Settings, go to LLM Provider, and select Add Custom Provider. Enter a display name, a lowercase model prefix such as vllm, the endpoint's base URL ending in /v1, and the API key if the server requires one. The custom provider documentation describes the setup.
3. Keep private endpoints private
If the GPU server runs in a private network, start the Connic Bridge inside that network and select it in the provider's Route via Bridge dropdown. Connic then tunnels every model request through the bridge, without inbound firewall rules or a public endpoint. The Bridge documentation and the Bridge overview describe the setup.
4. Reference Kolibri-1 in the agent
The model ID combines the provider prefix with the model name that vLLM serves, which is the Hugging Face repository name by default:
version: "1.0"
name: support-agent-de
model: vllm/Aleph-Alpha/Kolibri-1
description: "Answers German-language support requests"
fallback_model: connic/qwen3.6-35b-a3b
reasoning_effort: low
context_compression:
enabled: true
max_prompt_tokens: 200000
system_prompt: |
Du beantwortest Supportanfragen unserer Kunden auf Deutsch...Connic passes reasoning_effort through to Kolibri-1's chat template. Without it, every model call runs at high effort, so low or medium suits routine steps and high suits multi-step reasoning. With fallback_model, a failed or timed-out request moves to a Connic-managed model, so a GPU outage does not stop the agent. Context compression helps keep prompts below the recommended 262,144-token window.
vllm/ prefix is a provider alias set in the project, not an official Aleph Alpha model ID. Switching an agent between Kolibri-1 and a managed model such as connic/qwen3.6-35b-a3b changes only the model line.5. Test before rollout
Benchmarks do not show how Kolibri-1 handles a specific agent's tools and documents. A Connic A/B test does: the current agent stays as the control, a variant file named support-agent-de-test-kolibri.yaml changes the model, and part of the live traffic goes to the variant. Judges score both versions on the same criteria, and the test pauses a variant that breaches its failure-rate or quality limits. For tool calls with real consequences, approvals add the human review that the model card recommends.
Connect your Kolibri-1 endpoint as a custom provider and test it against EU-hosted managed models with A/B tests, judges, and traces.
Get started freeSelf-host Kolibri-1 or use a managed model?
Self-hosting Kolibri-1 makes sense when your German-language data has to stay on infrastructure you control and your team already operates H100- or H200-class GPUs. You then own capacity planning, updates, monitoring, and on-call for the inference server.
connic/qwen3.6-35b-a3b, connic/gemma-4-26b-a4b-it, connic/mistral-small-2603, and connic/qwen3.8-27b. The model catalog lists current prices and context limits.Because the model is a single line of agent YAML, you can start on a managed model today and A/B test a self-hosted Kolibri-1 endpoint once your hardware is in place. Judges and traces then show whether Kolibri-1's results on your German-language tasks justify running the inference server.