Mistral Large 4 is a 1-trillion-parameter multimodal language model that Mistral AI released as a public preview on October 6, 2026. Mistral's documentation lists 1.05 trillion total parameters, 52 billion active per token, and a 1.6-billion-parameter vision encoder. The preview is available through Mistral's API now, and Mistral says the weights will follow by the end of October. The release came three days after Aleph Alpha's Kolibri-1, another new model from a European developer. This guide covers the model's specifications, benchmarks, pricing, and weights timeline, compares it with Kolibri-1, and shows how to run it behind production AI agents.
Mistral Large 4 at a glance
Mistral Large 4 is a mixture-of-experts model that handles instruction following, reasoning, and agent tasks in one model. The specifications below come from the release announcement and the model page in Mistral's documentation.
| Specification | Mistral Large 4 |
|---|---|
| Developer | Mistral AI, France |
| Release | Public preview, October 6, 2026 |
| Parameters | 1.05T total, 52B active per token, plus a 1.6B vision encoder |
| Architecture | Granular mixture of experts, one model for instruction following and reasoning |
| Input and output | Text and images in, text out |
| Languages | More than 160 in the training data, including every official EU language |
| Context window | 1M tokens in Mistral's documentation; 524k in preview listings |
| Reasoning | Adjustable through the reasoning_effort parameter |
| API features | Function calling, structured outputs, document Q&A, batching, agents |
| API model IDs | mistral-large-4, mistral-large-4-0 |
| Weights | Promised by the end of October 2026 |
| License | Not yet announced |
| Training | From scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European data centers |
The sources differ on some figures. The announcement describes 1 trillion parameters with 49 billion active, while the documentation lists 1.05 trillion and 52 billion. The documentation gives a 1-million-token context window, but Artificial Analysis and OpenRouter list 524k tokens for the preview. Mistral also says the reinforcement learning run behind the preview is still in progress and expects further improvements, so results may change before the weights ship.
Mistral Large 4 benchmarks
Mistral presents most results as charts. The table collects the scores the announcement states in its text. Mistral measured most of them itself and chose the comparison models; the AA-Briefcase score comes from Artificial Analysis.
| Benchmark | Mistral Large 4 | What it measures |
|---|---|---|
| DeepSWE v1.1 | 61.7% | Software engineering tasks |
| SWE-Atlas-QnA | 59.4% | Repository understanding |
| Terminal-Bench 4 | 28.3% | Command-line workflows |
| Coding Agent Index | 49.8% | Ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max |
| AutomationBench | 59.9% | 657 business workflows; ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro |
| AA-Briefcase | 1,393 Elo | Long-horizon knowledge work; ahead of DeepSeek V4 Pro |
| Cybench | 93% | 40 challenges from security competitions |
| Dense 200 | 42% | Visual grounding; GPT-6 Astra scores 41% |
| B3 AI Security Benchmark | 93.3% | Share of attacks resisted |
Scores stated in Mistral's release announcement, October 6, 2026.
Coding and agent results
On AutomationBench, which covers 657 business workflows across apps such as Gmail, Google Sheets, Slack, and Salesforce, Mistral reports a score of 59.9%. Its combined Coding Agent Index of 49.8% places it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind evaluation of coding quality that Mistral ran with Surge AI, professional annotators rated the preview 3.74 out of 5, second of five models: behind Claude Opus 5 (4.22) and ahead of GLM-5.3 (3.60), Kimi K3 (3.59), and GLM-5.2 (3.40).
Cybersecurity
Mistral gives cybersecurity the most space in its announcement. On the Artificial Analysis Cyber Index, Mistral reports a top-five ranking worldwide and the lead among open-weight models developed outside China. On the index test that asks a model to reproduce a real vulnerability in open-source software and then patch it, Mistral Large 4 scores 82%, which Mistral calls the highest of any model. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score near zero on that test because they refuse the task. Mistral also reports a higher refusal rate on malicious cyber prompts than any other open model it tested. Until the weights ship, vetted partners and state authorities red-team a version with reduced moderation.
Images and knowledge work
Mistral describes the image capabilities as a step change from its earlier models, with reasoning over documents, charts, technical drawings, and satellite imagery. On visual grounding, it reports 42% on Dense 200 against 41% for GPT-6 Astra. In third-party evaluations by vals.ai on legal and financial tasks, the preview exceeds GPT-6 Astra, and on Harvey's Legal Agent Benchmark it outperforms all open-source models, according to Mistral.
Independent results
Artificial Analysis scores the preview 38 on its Intelligence Index, with reasoning, up from 9 for the non-reasoning Mistral Large 3. Open-weight models from China still score higher, including GLM-5.3 at 45 and Kimi K3 at 44, both in their maximum reasoning settings. Artificial Analysis lists the preview as a proprietary model because the weights are not yet public. It measured 116 output tokens per second and noted that the model is verbose: it generated 200 million output tokens during the index evaluation, against a median of 81 million for reasoning models in a similar price tier.
Pricing and API access
The preview runs on Mistral's API under the model IDs mistral-large-4 and mistral-large-4-0. Mistral lists the prices below per million tokens. Its documentation currently shows them crossed out at half that rate, without saying how long the reduced price applies.
| Tokens | List price | Current price in the docs |
|---|---|---|
| Input | $1.36 | $0.68 |
| Cached input | $0.14 | $0.07 |
| Output | $4.18 | $2.09 |
Prices per million tokens in US dollars, from Mistral's announcement and documentation on October 6, 2026.
Mistral offers batch processing at a 50% discount. For EU data location requirements, its regional inference endpoint api.eu.mistral.ai processes requests in data centers in EU and EFTA countries at 1.1 times the standard list price, while the global endpoint makes no commitment on processing location. Regional endpoints serve only the models hosted in their region and support function calling but not Mistral's agents, batch, or files features, and account and billing data may still be handled outside the region. Artificial Analysis found the preview verbose, so output tokens deserve attention in any cost estimate.
Weights and self-hosting
Mistral plans to publish the weights by the end of October, together with details on the architecture, more benchmarks, and its post-training method. It has not named a license yet. Its predecessor, Mistral Large 3, is available under Apache 2.0.
Mistral says the model will run on private cloud or on premises, but the size sets a high bar. At 8 bits per parameter, 1.05 trillion parameters take roughly 1 TB of memory, and about 2 TB at 16 bits, before KV cache and runtime overhead. That is more than an eight-GPU H100 server holds (8 × 80 GB = 640 GB). Mistral has not published hardware requirements yet, so for most teams the API remains the practical way to use the model.
Mistral Large 4 vs Kolibri-1: two European releases in one week
Aleph Alpha released Kolibri-1 three days before Mistral Large 4. The two releases follow opposite strategies: Mistral built a 1-trillion-parameter multilingual model with image input, while Aleph Alpha built a compact German-English model that fits on a single large GPU.
| Attribute | Mistral Large 4 | Kolibri-1 |
|---|---|---|
| Developer | Mistral AI, France | Aleph Alpha, Germany |
| Release | Public preview, October 6, 2026 | Weights released October 3, 2026 |
| Total parameters | 1.05T | 78.1B |
| Active per token | 52B | 3.46B |
| Input | Text and images | Text |
| Languages | More than 160 in the training data | German and English |
| Context window | 1M (Mistral docs), 524k (preview listings) | 262,144 native, validated to 1,048,576 |
| Weights | Promised by the end of October | Available on Hugging Face |
| License | Not yet announced | Apache 2.0 for weights and configuration files |
| Weight memory | About 1 TB at 8 bits (estimate) | About 78 GB (FP8) |
| Access today | Mistral API and API routers | Self-hosting with vLLM |
The benchmark results do not compare directly, because the two companies report different test suites. In practice, Mistral Large 4 suits multilingual products, document and image workloads, and teams that want a large multimodal model through an API. Kolibri-1 suits German-language workloads that must run on infrastructure a team controls, at a fraction of the hardware cost.
What Mistral Large 4 means for Europe
Mistral trained the model from scratch in its own European data centers and serves the preview on the same infrastructure. Mistral says the model will be available in several regions worldwide, including a European deployment that Mistral operates end to end, independently of other digital service providers and under European law. According to Mistral, the training data covers every official language of the EU, which matters for products that serve customers in several EU countries.
Mistral calls the model the first milestone on the roadmap funded by its €3 billion Series D, which it describes as the largest equity round ever raised by a European technology company. Our overview of EU-hosted AI models compares Mistral's EU endpoint with self-hosted open-weight models and other EU hosting options.
How to use Mistral Large 4 with Connic
Connic runs AI agents in production, with deployments, connectors, guardrails, approvals, judges, and traces built in. The Connic-managed catalog already includes Mistral Large 3 as connic/mistral-large-2512 and Mistral Medium 3.5, and Mistral Large 4 is planned for it. Until then, agents use Mistral Large 4 through Mistral's API as a custom OpenAI-compatible provider, which takes three steps. Projects that already use OpenRouter with their own key can reference openrouter/mistralai/mistral-large-4-0 instead.
1. Add Mistral as a custom provider
Create an API key in Mistral Studio. In Connic, open Project Settings, go to LLM Provider, and select Add Custom Provider. Enter a display name, the model prefix mistral, the base URL https://api.mistral.ai/v1, and the API key. For inference in EU and EFTA data centers, use https://api.eu.mistral.ai/v1 instead and check that the model is available there. The custom provider documentation describes the setup.
2. Reference the model in the agent
The model ID combines the provider prefix with Mistral's model ID:
version: "1.0"
name: contract-review
model: mistral/mistral-large-4
description: "Reviews supplier contracts and flags risky clauses"
fallback_model: connic/mistral-large-2512
system_prompt: |
You review supplier contracts...With fallback_model, a failed or timed-out request moves to the Connic-managed Mistral Large 3, so a preview outage does not stop the agent.
mistral/ prefix is a provider alias set in the project. Once Mistral Large 4 is in the Connic-managed catalog, switching changes only the model line, and the provider key is no longer needed.3. Test before rollout
Preview results can change from week to week, and the model writes long outputs. A Connic A/B test keeps the current agent as the control and sends part of the live traffic to a variant file such as contract-review-test-large-4.yaml. Judges score both versions on the same criteria, the runs table shows token counts per run, and the test pauses a variant that breaches its failure-rate or quality limits.
Connect Mistral Large 4 as a custom provider today and compare it with EU-hosted managed models using A/B tests, judges, and traces.
Get started free