Skip to main content
Connic
Back to BlogIndustry Insights

Aleph Alpha Kolibri-1: EU Open-Weight Model, Benchmarks & Setup

Aleph Alpha's Kolibri-1 is a 78B open-weight German-English model trained in Europe; this guide covers its benchmarks, hardware, license, and agent setup.

October 4, 202613 min readAuthor: Connic Research Team

Kolibri-1 is an open-weight language model that Aleph Alpha released on October 3, 2026. It has 78.1 billion parameters, activates 3.46 billion of them per token, and is built for German and English. It supports adjustable reasoning and tool calling, and its weights are available on Hugging Face under the Apache 2.0 license. Aleph Alpha trained the model on infrastructure in Germany and Finland. This guide covers the model's specifications, license, benchmark results, hardware requirements, and community quantizations, and shows how to run a self-hosted Kolibri-1 endpoint behind production AI agents.

Kolibri-1 at a glance

Kolibri-1 is a mixture-of-experts model: each layer holds many expert networks, and a router sends each token to a small subset of them. Compute per token stays close to that of a 3.5B model, while the full network stores 78.1B parameters. The specifications below come from the official Kolibri-1 model card and Aleph Alpha's release announcement.

Kolibri-1 technical specifications
SpecificationKolibri-1
DeveloperAleph Alpha, Heidelberg, Germany
Release dateOctober 3, 2026
Parameters78.1B total, 3.46B active per token
ArchitectureMixture of experts, 50 layers, 384 routed experts plus 1 shared expert per layer, 6 routed experts per token
AttentionSliding-window attention over 512 tokens in four of every five layers, full attention in the fifth
Context window262,144 tokens native, validated up to 1,048,576
LanguagesGerman and English
ReasoningAdjustable effort: none, low, medium, high
Tool callingYes, through the official vLLM parser
WeightsFP8 (about 78 GB) and BF16 (about 156 GB)
Knowledge cutoffJune 18, 2026
Pre-training data20 trillion tokens, 21.3% of them German (about 4.3 trillion), per the release announcement
Pre-training compute768 NVIDIA B200 GPUs for 21 days
Training locationInfrastructure in Germany and Finland
LicenseApache 2.0 for the published weights and configuration files

Aleph Alpha also trained a new tokenizer that follows German word structure, including compound nouns. The model card reports 4.7 bytes per token on German text, a higher compression than leading models with larger vocabularies, while English stays at 4.2 bytes per token. Fewer tokens for the same German text mean lower latency and serving cost per request.

The Apache 2.0 license allows commercial use, modification, and redistribution of the weights. Its scope is defined narrowly: according to the model card, the rights apply to the weights and configuration files in the Hugging Face repository and do not extend to code, model architecture, parameter settings, or training methods. The serving plugin, aleph-alpha-inference, is published separately under Apache 2.0.

Kolibri-1 benchmarks

Aleph Alpha compared the post-trained model with 13 other models, mostly using its open-source eval-framework with the same setup for every model and Kolibri-1 at reasoning effort high. The comparison groups mixture-of-experts models with similar active parameter counts and lists dense models, which activate several times more parameters per token, for reference. All figures below are Aleph Alpha's own measurements, not independent results.

Compared with similar open models

The table shows Kolibri-1 next to three mixture-of-experts models of similar active size and Qwen3.8 27B, the higher-scoring of the two dense models in Aleph Alpha's comparison.

Selected post-training results from the Kolibri-1 model card, October 2026
BenchmarkKolibri-1Qwen3.6 35B-A3BGemma 4 26B-A4BMistral Small 4 119B-A6BQwen3.8 27B (dense)
Overall, English75.571.471.963.180.2
Overall, German70.867.366.361.479.9
Agentic average (EN)63.462.154.640.766.7
Tau3-Bench Banking38.110.616.05.750.0
BFCL v4 multi-turn47.558.161.440.455.5
Agentic Wiki QA (DE)69.458.866.265.473.5
Industry RAG average (DE)67.565.849.453.480.2
LiveCodeBench v685.982.582.371.293.8
AA-Omniscience non-hallucination rate44.056.714.334.767.3

Selected post-training results from the Kolibri-1 model card (FP8 release), October 2026. Kolibri-1 ran at reasoning effort high; the other models used their documented settings. Bold marks the best mixture-of-experts score in each row; the shaded column is the dense reference model.

Kolibri-1 posts the highest overall English and German scores of every mixture-of-experts model in the comparison, but the margins are narrow. Qwen3.5 35B-A3B reaches 74.7 in English and 69.8 in German, and GPT-OSS 120B scores 70.2 in German. Qwen3.8 27B leads Kolibri-1 by about 5 points in English and 9 in German, while activating almost eight times as many parameters per token.

Agent and tool-calling results

Kolibri-1's English agentic average of 63.4 ties Qwen3.5 35B-A3B for the best mixture-of-experts result. Its clearest lead is on Tau3-Bench Banking, a simulated banking customer-service task, with 38.1 against 16.0 or less for every other mixture-of-experts model. It reaches 76.7 on Tau2-Bench Airline, the best mixture-of-experts score there, and 69.4 on Aleph Alpha's German agentic Wiki QA task, also the best in its group.

Multi-turn function calling is the weak spot. On the multi-turn split of the Berkeley Function Calling Leaderboard (BFCL v4), Kolibri-1 scores 47.5, below 9 of the 11 other mixture-of-experts models, and its overall BFCL v4 score of 61.4 trails Qwen3.6 35B-A3B (67.2) and Gemma 4 26B-A4B (68.2). On TerminalBench 2.1, a command-line agent benchmark, it scores 27.7 against 39.7 for Qwen3.5 35B-A3B. Agents that chain many dependent tool calls within one conversation should be tested on exactly that pattern before rollout.

Grounding and hallucinations

Aleph Alpha trained Kolibri-1 with abstention data so that it declines questions it cannot answer reliably. On the RGB negative-rejection test, which checks whether a model declines when the supplied documents contain no answer, it scores 85.6. Its AA-Omniscience non-hallucination rate of 44.0 trails Qwen3.6 35B-A3B (56.7) but is far ahead of Gemma 4 26B-A4B (14.3). Correcting false information in supplied documents is weaker: on the RGB fact-check test it scores 34.0, while the mixture-of-experts models from other developers range from 58.0 to 90.0. Retrieval agents built on Kolibri-1 therefore benefit from checks that compare answers against their sources.

Kolibri-1's final long-context training phase used 262,144-token sequences, which is its native context length. Positional encoding sits only in the sliding-window layers, so the context can grow further without position scaling, and Aleph Alpha validated quality and serving efficiency up to 1,048,576 tokens. The model card still recommends at most 262,144 tokens for latency- or throughput-sensitive deployments and for complex tasks. The default configuration serves 262,144 tokens; longer contexts need the additional vLLM flags --max-model-len 1048576 and --hf-overrides '{"max_position_embeddings": 1048576}'.

Long-context results for the base model are consistent with that recommendation. Its RULER average falls from 86.9 at 4k tokens to 69.8 at 256k and 63.2 at 1M. At 1M tokens it scores higher than the base versions of Nemotron 3 Nano 30B-A3B (58.5) and Qwen3.5 35B-A3B (57.5), but both of them score higher at every length up to 512k.

Agents rarely need a million tokens in a single request. A long-running agent accumulates tool results and conversation history, and keeping each request under the recommended window protects both latency and answer quality. Connic's context compression shortens the history once prompts reach a configured token limit; the guide to context compression explains the trade-offs.

Hardware requirements and serving

The sparse design reduces compute per token, but all 78.1B parameters must stay in GPU memory. Aleph Alpha lists the following configurations; the KV cache and runtime overhead come on top of the weights.

Kolibri-1 memory footprint and GPU configurations
WeightsWeight memoryMinimumRecommended
FP8 (Kolibri-1)About 78 GB2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B3002× H100 SXM5, 2× H200, 1× B200 or 1× B300
BF16 (Kolibri-1-BF16)About 156 GB4× A100 80 GB, 4× H100 SXM5, 2× H200, 1× B200 or 1× B3004× H100 SXM5, 2× H200, 2× B200 or 1× B300

Serving requires Aleph Alpha's aleph-alpha-inference package, which adds Kolibri-1 support to vLLM and installs the vLLM version it supports. A prebuilt container image is available as well. The official command enables reasoning and tool calling:

Serve Kolibri-1 with vLLM
pip install 'aleph-alpha-inference>=1'

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

Aleph Alpha recommends temperature 1.0, top_p 0.97, and top_k 128, and the repository ships these values in its generation config. Reasoning effort is set per request through the chat template, with the levels none, low, medium, and high. When a request sets no effort, the template uses high, which adds reasoning tokens and latency to every call. Tool calls use a Hermes-style format that the kolibri1 parser converts into structured tool calls in the OpenAI-compatible response, and tool calling works together with reasoning.

Quantized versions: GGUF, MLX, and NVFP4

Aleph Alpha publishes only the FP8 and BF16 checkpoints. Within a day of the release, community members uploaded more than a dozen GGUF, MLX, and NVFP4 conversions to Hugging Face. As of October 4, 2026, the GGUF and MLX files need patched or forked runtimes. A selection:

Selected unofficial Kolibri-1 quantizations on Hugging Face as of October 4, 2026
FormatRepositorySizeRuntimeNotes
GGUF Q3_K_SEliasfpv28/Kolibri-1-Q3_K_S-GGUF33.9 GBPatched llama.cppStarted locally at a 4,096-token context; reasoning and tool calling not validated
GGUF Q4_K_M, Q8_0Hob-forge/Kolibri-1-GGUF47.5 GB, 83.1 GBPatched llama.cppArchitecture patch checked against a reference implementation
MLX 3-bithere-be-dragons-ai/Kolibri-1-MLX-3bit33 GiBmlx-vlm forkSized for Macs with 48 GB unified memory
MLX 2-bitvelaia/Kolibri-1-MLX-2bit24 GiBmlx-lm with a bundled launcherMacs with 36 GB or more
MLX 4-bitvelaia/Kolibri-1-MLX-4bit41 GiBmlx-lm with a bundled launcherMacs with 64 GB or more
MLX 6-bitaudreyt/Kolibri-1-MLX-6bit59.16 GiBmlx-lm forkMacs with 96 GB or more
NVFP4 W4A16audreyt/Kolibri-1-NVFP4-W4A1644.14 GiBvLLMFull generation not yet validated

Stock llama.cpp, Ollama, and LM Studio cannot load the GGUF files until they support the kolibri1 architecture. On October 4, a llama.cpp feature request and an mlx-lm pull request for Kolibri-1 were still open. The NVFP4 build targets vLLM with Aleph Alpha's plugin but has not completed a generation test yet. Most conversions start from the FP8 checkpoint, and their authors report smoke tests and small perplexity samples rather than benchmark results.

Unofficial conversions
None of these files comes from Aleph Alpha, and lower bit widths change model quality. They suit local experiments on a workstation or Mac. For production agents, the official FP8 weights on vLLM remain the reference setup.

What Kolibri-1 means for Europe

A German company developed Kolibri-1, and according to the release announcement it was trained in Germany and Finland under European and German law. The Apache 2.0 weights run in any EU data center or on premises, so prompts and documents stay on infrastructure the operator controls. Teams comparing hosting options can read our overview of EU-hosted AI models.

The release documentation also addresses EU rules for general-purpose AI models. The model card states that Aleph Alpha is a signatory of the EU General-Purpose AI Code of Practice and names a contact for rightsholders; its responsible-use guidance asks users not to apply the model to practices prohibited under Article 5 of the EU AI Act. It describes the intended deployment as assistants and agentic workflows in which a person reviews the output before it is acted on, rather than systems that act unreviewed. Obligations for a deployed agent still depend on its use case; the guide to EU AI Act compliance for AI agents covers that assessment.

German and English are the only supported languages. That makes Kolibri-1 a candidate for German-language workloads in Germany, Austria, and Switzerland and for European companies serving German-speaking customers, but not for products that need the EU's other languages. Aleph Alpha positions the model for public administration, aerospace, manufacturing, automotive, and semiconductor companies.

How to use Kolibri-1 with Connic

Connic runs AI agents in production with deployments, connectors, guardrails, approvals, judges, and traces built in. Kolibri-1 is not part of the Connic-managed model catalog, but a self-hosted Kolibri-1 endpoint connects as a custom OpenAI-compatible provider and then works with all of these features like any other model. The setup takes five steps.

1. Serve the model

Run the official vLLM command from the previous section on a GPU server that meets the hardware requirements. If the endpoint is reachable from outside its network, start vLLM with --api-key so that only clients with the key can call it.

2. Add a custom provider

In Connic, open Project Settings, go to LLM Provider, and select Add Custom Provider. Enter a display name, a lowercase model prefix such as vllm, the endpoint's base URL ending in /v1, and the API key if the server requires one. The custom provider documentation describes the setup.

3. Keep private endpoints private

If the GPU server runs in a private network, start the Connic Bridge inside that network and select it in the provider's Route via Bridge dropdown. Connic then tunnels every model request through the bridge, without inbound firewall rules or a public endpoint. The Bridge documentation and the Bridge overview describe the setup.

4. Reference Kolibri-1 in the agent

The model ID combines the provider prefix with the model name that vLLM serves, which is the Hugging Face repository name by default:

agents/support-agent-de.yaml
version: "1.0"

name: support-agent-de
model: vllm/Aleph-Alpha/Kolibri-1
description: "Answers German-language support requests"
fallback_model: connic/qwen3.6-35b-a3b
reasoning_effort: low
context_compression:
  enabled: true
  max_prompt_tokens: 200000
system_prompt: |
  Du beantwortest Supportanfragen unserer Kunden auf Deutsch...

Connic passes reasoning_effort through to Kolibri-1's chat template. Without it, every model call runs at high effort, so low or medium suits routine steps and high suits multi-step reasoning. With fallback_model, a failed or timed-out request moves to a Connic-managed model, so a GPU outage does not stop the agent. Context compression helps keep prompts below the recommended 262,144-token window.

One line to switch models
The vllm/ prefix is a provider alias set in the project, not an official Aleph Alpha model ID. Switching an agent between Kolibri-1 and a managed model such as connic/qwen3.6-35b-a3b changes only the model line.

5. Test before rollout

Benchmarks do not show how Kolibri-1 handles a specific agent's tools and documents. A Connic A/B test does: the current agent stays as the control, a variant file named support-agent-de-test-kolibri.yaml changes the model, and part of the live traffic goes to the variant. Judges score both versions on the same criteria, and the test pauses a variant that breaches its failure-rate or quality limits. For tool calls with real consequences, approvals add the human review that the model card recommends.

Run Kolibri-1 behind production agents

Connect your Kolibri-1 endpoint as a custom provider and test it against EU-hosted managed models with A/B tests, judges, and traces.

Get started free

Self-host Kolibri-1 or use a managed model?

Self-hosting Kolibri-1 makes sense when your German-language data has to stay on infrastructure you control and your team already operates H100- or H200-class GPUs. You then own capacity planning, updates, monitoring, and on-call for the inference server.

Start without GPUs
All four comparison models from the benchmark table run as Connic-managed models with EU-only inference, no provider account, and token billing from project credit: connic/qwen3.6-35b-a3b, connic/gemma-4-26b-a4b-it, connic/mistral-small-2603, and connic/qwen3.8-27b. The model catalog lists current prices and context limits.

Because the model is a single line of agent YAML, you can start on a managed model today and A/B test a self-hosted Kolibri-1 endpoint once your hardware is in place. Judges and traces then show whether Kolibri-1's results on your German-language tasks justify running the inference server.

Frequently Asked Questions

Kolibri-1 is an open-weight mixture-of-experts language model from Aleph Alpha, released on October 3, 2026. It has 78.1 billion parameters, activates 3.46 billion per token, supports German and English, offers adjustable reasoning and tool calling, and is validated for contexts of up to 1,048,576 tokens. Its weights are available on Hugging Face under Apache 2.0.

Kolibri-1 is open-weight. Aleph Alpha publishes FP8 and BF16 weights under Apache 2.0, which allows commercial use, modification, and redistribution. According to the model card, the license covers the weights and configuration files in the repository and does not extend to code, model architecture, parameter settings, or training methods.

The FP8 weights need about 78 GB of GPU memory. Aleph Alpha lists 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200, or 1× B300 as minimum configurations. The BF16 weights need about 156 GB. KV cache and runtime overhead come on top of the weights.

Kolibri-1's native context is 262,144 tokens, and Aleph Alpha validated it up to 1,048,576 tokens. The model card recommends at most 262,144 tokens for latency- or throughput-sensitive deployments and complex tasks. Serving more than 262,144 tokens requires additional vLLM flags.

Not with stock tools as of October 4, 2026. Community GGUF files need a patched llama.cpp build, and Ollama and LM Studio cannot load them yet. Community MLX conversions run on Apple Silicon Macs through forked runtimes or a bundled launcher and need at least 36 GB of unified memory, more at higher bit widths. None of these conversions is official.

In Aleph Alpha's evaluation, Kolibri-1 scores 75.5 overall in English and 70.8 in German, ahead of Qwen3.6 35B-A3B (71.4 and 67.3) and Gemma 4 26B-A4B (71.9 and 66.3). The dense Qwen3.8 27B scores higher at 80.2 and 79.9 but activates almost eight times as many parameters per token. All results are Aleph Alpha's own.

Yes. The official vLLM setup enables Hermes-style tool calling through the kolibri1 parser, and tool calls work together with reasoning. In Aleph Alpha's benchmarks, Kolibri-1 is strong on simulated customer-service tasks such as Tau3-Bench Banking but weaker at multi-turn function calling on BFCL v4.

According to the model card, Aleph Alpha is a signatory of the EU General-Purpose AI Code of Practice, and its responsible-use guidance asks users not to apply Kolibri-1 to practices prohibited under Article 5 of the EU AI Act. Self-hosting keeps data on infrastructure the operator controls. Obligations for a deployed agent still depend on its use case and must be assessed for each system.

Yes. Serve Kolibri-1 with vLLM, add the endpoint as a custom OpenAI-compatible provider in Project Settings, and reference it in agent YAML, for example as vllm/Aleph-Alpha/Kolibri-1. Private endpoints connect through the Connic Bridge. Kolibri-1 is not part of the Connic-managed model catalog.

More from the Blog

Industry Insights

Who Owns EU AI Act Readiness? Evidence from 2,390 DPOs

A 2025 survey of 2,390 DPOs shows EU AI Act work landing with them; here is how product, engineering, legal, and security teams can share it.

September 28, 202610 min read
Industry Insights

AI Agent Platform SLA Checklist: What Enterprise Buyers Should Verify

Evaluate an AI agent platform SLA across uptime scope, dependencies, incident response, recovery, security evidence, remedies, and exit terms.

August 29, 202612 min read
Industry Insights

EU-Hosted AI Models in 2026: Providers, Dependence & Options

EU-hosted AI models compared by location, retention, operator, portability, legal exposure, and deployment model, using a 2026 Commission-requested study.

August 18, 202611 min read
Industry Insights

EU AI Gigafactories: What the €30B Plan Means for Enterprise AI

The EU's €30B Gigafactory plan could expand compute, while eligible AI SMEs can apply for AI Factory access through Playground, Fast Lane or Large Scale.

August 14, 202615 min read
Industry Insights

The OpenAI Hugging Face Hack: Guardrail Lessons for AI Agents

OpenAI models escaped a test sandbox and breached Hugging Face in July 2026. What the incident reveals about guardrails and how to secure production AI agents.

July 24, 20269 min read
Industry Insights

AI Agent Platforms With EU Data Residency: 2026 Shortlist

AI agent platforms compared by EU data residency, including where traces, storage, model calls, backups, subprocessors, and support access are processed.

July 6, 202612 min read
Industry Insights

How to Run AI Agents in the EU Without US Hyperscalers

Production AI agents in the EU without US hyperscalers: what EU-hosted must really mean, exposure under the US CLOUD Act, and a sovereignty checklist.

June 4, 20269 min read
Industry Insights

Best AI Agent Platforms for EU Enterprises in 2026

Ranked shortlist of AI agent platforms evaluated on EU data residency, self-hosting, MCP tool support, BYOK, EU AI Act readiness, and SLA terms. Updated September 2026.

May 19, 202616 min read
Industry Insights

AI Agent Deployment Platforms: 16 Vendors Compared (2026)

Sixteen AI agent deployment platforms compared by runtime boundary, language, hosting model, connector ownership, residency, and pricing.

April 19, 202615 min read