Skip to main content
Connic
Back to BlogIndustry Insights

Mistral Large 4: Europe's 1T Model, Benchmarks, Pricing & Setup

Mistral Large 4 is a 1T-parameter multimodal model trained in Europe; this guide covers its benchmarks, pricing, weights timeline, and setup for AI agents.

October 6, 202613 min readAuthor: Connic Research Team

Mistral Large 4 is a 1-trillion-parameter multimodal language model that Mistral AI released as a public preview on October 6, 2026. Mistral's documentation lists 1.05 trillion total parameters, 52 billion active per token, and a 1.6-billion-parameter vision encoder. The preview is available through Mistral's API now, and Mistral says the weights will follow by the end of October. The release came three days after Aleph Alpha's Kolibri-1, another new model from a European developer. This guide covers the model's specifications, benchmarks, pricing, and weights timeline, compares it with Kolibri-1, and shows how to run it behind production AI agents.

Mistral Large 4 at a glance

Mistral Large 4 is a mixture-of-experts model that handles instruction following, reasoning, and agent tasks in one model. The specifications below come from the release announcement and the model page in Mistral's documentation.

Mistral Large 4 technical specifications
SpecificationMistral Large 4
DeveloperMistral AI, France
ReleasePublic preview, October 6, 2026
Parameters1.05T total, 52B active per token, plus a 1.6B vision encoder
ArchitectureGranular mixture of experts, one model for instruction following and reasoning
Input and outputText and images in, text out
LanguagesMore than 160 in the training data, including every official EU language
Context window1M tokens in Mistral's documentation; 524k in preview listings
ReasoningAdjustable through the reasoning_effort parameter
API featuresFunction calling, structured outputs, document Q&A, batching, agents
API model IDsmistral-large-4, mistral-large-4-0
WeightsPromised by the end of October 2026
LicenseNot yet announced
TrainingFrom scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European data centers

The sources differ on some figures. The announcement describes 1 trillion parameters with 49 billion active, while the documentation lists 1.05 trillion and 52 billion. The documentation gives a 1-million-token context window, but Artificial Analysis and OpenRouter list 524k tokens for the preview. Mistral also says the reinforcement learning run behind the preview is still in progress and expects further improvements, so results may change before the weights ship.

Mistral Large 4 benchmarks

Mistral presents most results as charts. The table collects the scores the announcement states in its text. Mistral measured most of them itself and chose the comparison models; the AA-Briefcase score comes from Artificial Analysis.

Mistral Large 4 results stated in the release announcement, October 2026
BenchmarkMistral Large 4What it measures
DeepSWE v1.161.7%Software engineering tasks
SWE-Atlas-QnA59.4%Repository understanding
Terminal-Bench 428.3%Command-line workflows
Coding Agent Index49.8%Ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max
AutomationBench59.9%657 business workflows; ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro
AA-Briefcase1,393 EloLong-horizon knowledge work; ahead of DeepSeek V4 Pro
Cybench93%40 challenges from security competitions
Dense 20042%Visual grounding; GPT-6 Astra scores 41%
B3 AI Security Benchmark93.3%Share of attacks resisted

Scores stated in Mistral's release announcement, October 6, 2026.

Coding and agent results

On AutomationBench, which covers 657 business workflows across apps such as Gmail, Google Sheets, Slack, and Salesforce, Mistral reports a score of 59.9%. Its combined Coding Agent Index of 49.8% places it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind evaluation of coding quality that Mistral ran with Surge AI, professional annotators rated the preview 3.74 out of 5, second of five models: behind Claude Opus 5 (4.22) and ahead of GLM-5.3 (3.60), Kimi K3 (3.59), and GLM-5.2 (3.40).

Cybersecurity

Mistral gives cybersecurity the most space in its announcement. On the Artificial Analysis Cyber Index, Mistral reports a top-five ranking worldwide and the lead among open-weight models developed outside China. On the index test that asks a model to reproduce a real vulnerability in open-source software and then patch it, Mistral Large 4 scores 82%, which Mistral calls the highest of any model. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score near zero on that test because they refuse the task. Mistral also reports a higher refusal rate on malicious cyber prompts than any other open model it tested. Until the weights ship, vetted partners and state authorities red-team a version with reduced moderation.

Images and knowledge work

Mistral describes the image capabilities as a step change from its earlier models, with reasoning over documents, charts, technical drawings, and satellite imagery. On visual grounding, it reports 42% on Dense 200 against 41% for GPT-6 Astra. In third-party evaluations by vals.ai on legal and financial tasks, the preview exceeds GPT-6 Astra, and on Harvey's Legal Agent Benchmark it outperforms all open-source models, according to Mistral.

Independent results

Artificial Analysis scores the preview 38 on its Intelligence Index, with reasoning, up from 9 for the non-reasoning Mistral Large 3. Open-weight models from China still score higher, including GLM-5.3 at 45 and Kimi K3 at 44, both in their maximum reasoning settings. Artificial Analysis lists the preview as a proprietary model because the weights are not yet public. It measured 116 output tokens per second and noted that the model is verbose: it generated 200 million output tokens during the index evaluation, against a median of 81 million for reasoning models in a similar price tier.

Pricing and API access

The preview runs on Mistral's API under the model IDs mistral-large-4 and mistral-large-4-0. Mistral lists the prices below per million tokens. Its documentation currently shows them crossed out at half that rate, without saying how long the reduced price applies.

Mistral Large 4 API prices per million tokens, October 6, 2026
TokensList priceCurrent price in the docs
Input$1.36$0.68
Cached input$0.14$0.07
Output$4.18$2.09

Prices per million tokens in US dollars, from Mistral's announcement and documentation on October 6, 2026.

Mistral offers batch processing at a 50% discount. For EU data location requirements, its regional inference endpoint api.eu.mistral.ai processes requests in data centers in EU and EFTA countries at 1.1 times the standard list price, while the global endpoint makes no commitment on processing location. Regional endpoints serve only the models hosted in their region and support function calling but not Mistral's agents, batch, or files features, and account and billing data may still be handled outside the region. Artificial Analysis found the preview verbose, so output tokens deserve attention in any cost estimate.

Weights and self-hosting

Mistral plans to publish the weights by the end of October, together with details on the architecture, more benchmarks, and its post-training method. It has not named a license yet. Its predecessor, Mistral Large 3, is available under Apache 2.0.

Mistral says the model will run on private cloud or on premises, but the size sets a high bar. At 8 bits per parameter, 1.05 trillion parameters take roughly 1 TB of memory, and about 2 TB at 16 bits, before KV cache and runtime overhead. That is more than an eight-GPU H100 server holds (8 × 80 GB = 640 GB). Mistral has not published hardware requirements yet, so for most teams the API remains the practical way to use the model.

Mistral Large 4 vs Kolibri-1: two European releases in one week

Aleph Alpha released Kolibri-1 three days before Mistral Large 4. The two releases follow opposite strategies: Mistral built a 1-trillion-parameter multilingual model with image input, while Aleph Alpha built a compact German-English model that fits on a single large GPU.

Mistral Large 4 and Kolibri-1 compared, October 2026
AttributeMistral Large 4Kolibri-1
DeveloperMistral AI, FranceAleph Alpha, Germany
ReleasePublic preview, October 6, 2026Weights released October 3, 2026
Total parameters1.05T78.1B
Active per token52B3.46B
InputText and imagesText
LanguagesMore than 160 in the training dataGerman and English
Context window1M (Mistral docs), 524k (preview listings)262,144 native, validated to 1,048,576
WeightsPromised by the end of OctoberAvailable on Hugging Face
LicenseNot yet announcedApache 2.0 for weights and configuration files
Weight memoryAbout 1 TB at 8 bits (estimate)About 78 GB (FP8)
Access todayMistral API and API routersSelf-hosting with vLLM

The benchmark results do not compare directly, because the two companies report different test suites. In practice, Mistral Large 4 suits multilingual products, document and image workloads, and teams that want a large multimodal model through an API. Kolibri-1 suits German-language workloads that must run on infrastructure a team controls, at a fraction of the hardware cost.

What Mistral Large 4 means for Europe

Mistral trained the model from scratch in its own European data centers and serves the preview on the same infrastructure. Mistral says the model will be available in several regions worldwide, including a European deployment that Mistral operates end to end, independently of other digital service providers and under European law. According to Mistral, the training data covers every official language of the EU, which matters for products that serve customers in several EU countries.

Mistral calls the model the first milestone on the roadmap funded by its €3 billion Series D, which it describes as the largest equity round ever raised by a European technology company. Our overview of EU-hosted AI models compares Mistral's EU endpoint with self-hosted open-weight models and other EU hosting options.

How to use Mistral Large 4 with Connic

Connic runs AI agents in production, with deployments, connectors, guardrails, approvals, judges, and traces built in. The Connic-managed catalog already includes Mistral Large 3 as connic/mistral-large-2512 and Mistral Medium 3.5, and Mistral Large 4 is planned for it. Until then, agents use Mistral Large 4 through Mistral's API as a custom OpenAI-compatible provider, which takes three steps. Projects that already use OpenRouter with their own key can reference openrouter/mistralai/mistral-large-4-0 instead.

1. Add Mistral as a custom provider

Create an API key in Mistral Studio. In Connic, open Project Settings, go to LLM Provider, and select Add Custom Provider. Enter a display name, the model prefix mistral, the base URL https://api.mistral.ai/v1, and the API key. For inference in EU and EFTA data centers, use https://api.eu.mistral.ai/v1 instead and check that the model is available there. The custom provider documentation describes the setup.

2. Reference the model in the agent

The model ID combines the provider prefix with Mistral's model ID:

agents/contract-review.yaml
version: "1.0"

name: contract-review
model: mistral/mistral-large-4
description: "Reviews supplier contracts and flags risky clauses"
fallback_model: connic/mistral-large-2512
system_prompt: |
  You review supplier contracts...

With fallback_model, a failed or timed-out request moves to the Connic-managed Mistral Large 3, so a preview outage does not stop the agent.

One line to switch models
The mistral/ prefix is a provider alias set in the project. Once Mistral Large 4 is in the Connic-managed catalog, switching changes only the model line, and the provider key is no longer needed.

3. Test before rollout

Preview results can change from week to week, and the model writes long outputs. A Connic A/B test keeps the current agent as the control and sends part of the live traffic to a variant file such as contract-review-test-large-4.yaml. Judges score both versions on the same criteria, the runs table shows token counts per run, and the test pauses a variant that breaches its failure-rate or quality limits.

Preview status
Mistral Large 4 is a public preview. Mistral expects further improvements from training that is still running, and the weights and license are not yet available. Plan production use with a fallback model and re-run evaluations when Mistral updates the model.
Run Mistral Large 4 behind production agents

Connect Mistral Large 4 as a custom provider today and compare it with EU-hosted managed models using A/B tests, judges, and traces.

Get started free

Frequently Asked Questions

Mistral Large 4 is a multimodal mixture-of-experts language model from Mistral AI, released as a public preview on October 6, 2026. Mistral's documentation lists 1.05 trillion total and 52 billion active parameters plus a 1.6-billion-parameter vision encoder. It accepts text and images, was trained on more than 160 languages, and is available through Mistral's API.

Not at the moment. Mistral promises the weights by the end of October 2026 but has not announced a license, so it is not yet clear whether the terms will be open source. Until then, the preview is available only through APIs, and Artificial Analysis lists it as a proprietary model. Its predecessor, Mistral Large 3, is available under Apache 2.0.

Mistral lists $1.36 per million input tokens, $0.14 per million cached input tokens, and $4.18 per million output tokens. On October 6, 2026, Mistral's documentation showed these prices reduced by half. Batch processing costs 50% less, and Mistral's EU regional endpoint costs 1.1 times the standard list price.

Mistral's documentation lists a context window of 1 million tokens. Artificial Analysis and OpenRouter list 524k tokens for the preview API.

Only after Mistral releases the weights, which it plans by the end of October 2026. With 1.05 trillion parameters, the weights take roughly 1 TB of memory at 8 bits per parameter, before KV cache and runtime overhead, so self-hosting requires multi-GPU servers. Mistral has not published hardware requirements yet.

Mistral Large 4 has 1.05 trillion parameters, accepts images, and covers more than 160 languages, but its weights are not yet released. Kolibri-1 from Aleph Alpha has 78.1 billion parameters, supports German and English only, and its weights are available under Apache 2.0 and fit in about 78 GB. Their benchmarks use different test suites and do not compare directly.

Yes. Mistral's API supports function calling and structured outputs for Mistral Large 4. Mistral reports 59.9% on AutomationBench, which covers 657 business workflows across apps such as Gmail, Google Sheets, Slack, and Salesforce.

Mistral says it trained the model in its own European data centers and serves the preview there, but its documentation makes no location commitment for the global endpoint. The EU regional endpoint processes inference requests in EU and EFTA data centers for 1.1 times the standard list price, if the model is available there; account and billing data may be handled outside the region.

Yes. Mistral's API connects as a custom OpenAI-compatible provider in Project Settings with the prefix mistral, and agent YAML then references mistral/mistral-large-4. Mistral Large 4 is planned for the Connic-managed catalog, which already includes Mistral Large 3 and Mistral Medium 3.5.

More from the Blog

Industry Insights

Aleph Alpha Kolibri-1: EU Open-Weight Model, Benchmarks & Setup

Aleph Alpha's Kolibri-1 is a 78B open-weight German-English model trained in Europe; this guide covers its benchmarks, hardware, license, and agent setup.

October 4, 202613 min read
Industry Insights

The OpenAI Hugging Face Hack: Guardrail Lessons for AI Agents

OpenAI models escaped a test sandbox and breached Hugging Face in July 2026. What the incident reveals about guardrails and how to secure production AI agents.

July 24, 20269 min read
Industry Insights

What Is an MCP Connector? A Practical Definition

An MCP connector links an AI app to external tools and data over the Model Context Protocol. Learn how it works and when it beats a custom API integration.

July 8, 20268 min read
Industry Insights

AI Agent Platforms With EU Data Residency: 2026 Shortlist

AI agent platforms compared by EU data residency, including where traces, storage, model calls, backups, subprocessors, and support access are processed.

July 6, 202612 min read
Industry Insights

Pre-built AI Agent Connectors: Platforms Compared (2026)

Pre-built AI agent connectors compared by platform, supported modes, documented recovery behavior, official sources, and production trade-offs.

June 16, 202615 min read
Industry Insights

The Real Cost of Assembling an AI Agent Stack

The real cost of assembling an AI agent stack comes from the integration and maintenance tax between tools. The comparison explains when an integrated platform makes sense.

June 9, 202610 min read
Industry Insights

How to Run AI Agents in the EU Without US Hyperscalers

Production AI agents in the EU without US hyperscalers: what EU-hosted must really mean, exposure under the US CLOUD Act, and a sovereignty checklist.

June 4, 20269 min read
Industry Insights

AI Agent Deployment Platforms: 16 Vendors Compared (2026)

Sixteen AI agent deployment platforms compared by runtime boundary, language, hosting model, connector ownership, residency, and pricing.

April 19, 202615 min read
Industry Insights

EU AI Act Enforcement: Who Investigates and What Evidence to Keep

The AI Office, national authorities, and EDPS divide EU AI Act enforcement by system and provider; teams should keep scoped governance and runtime evidence.

April 13, 202614 min read