# Mistral Large 4

> Source: https://aiwiki.ai/wiki/mistral_large_4
> Updated: 2026-10-06
> Fact-checked: 2026-10-06
> Categories: AI Models, Large Language Models, Mixture of Experts, Multimodal AI
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Mistral Large 4." aiwiki.ai, 6 Oct 2026. https://aiwiki.ai/wiki/mistral_large_4
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

Mistral Large 4 is a general-purpose [multimodal model](https://aiwiki.ai/wiki/vision_language_model) developed by [Mistral AI](https://aiwiki.ai/wiki/mistral_ai). It entered public preview on October 6, 2026, with a [mixture-of-experts](https://aiwiki.ai/wiki/mixture_of_experts) architecture and API access. The initial release preceded the planned distribution of its weights.[1] It belongs to the [Mistral Large](https://aiwiki.ai/wiki/mistral_large) family, but its preview specifications and availability should not be confused with those of [Mistral Large 3](https://aiwiki.ai/wiki/mistral_large_3).

## Release and availability

Mistral called the model ML4 and "le Chonk", offered a Studio preview API, and announced weights for the end of October 2026.[2]

On October 7, the model card listed weights and their license as forthcoming, not an available checkpoint with published licensing terms.[3]

Public preview is a distinct stage in Mistral's model lifecycle. Mistral permits silent updates during this stage and does not guarantee that a preview will reach general availability. Results obtained from the preview can therefore describe a changing service rather than a permanently fixed checkpoint.[6]

## Architecture and specifications

The specifications below describe the October 2026 preview.[3]

| Property | Published value |
|---|---|
| Architecture | Granular [mixture of experts](https://aiwiki.ai/wiki/mixture_of_experts)[3] |
| Total parameters | 1.05 trillion in the model card[3] |
| Vision encoder | 1.6 billion parameters[3] |
| Active parameters | 52 billion in the English model card; 49 billion in the launch announcement[3][2] |
| Mistral [context window](https://aiwiki.ai/wiki/context_window) | 1 million tokens in the model card[3] |
| OpenRouter context limit | 512K tokens, displayed as 524K in its summary[4] |
| OpenRouter maximum output | 256K tokens[4] |
| Input and output | Text and image input; text output on OpenRouter[4] |

The English card lists 52B active parameters; the announcement states 49B. These are conflicting published figures, not a reconciled count.[3][2]

OpenRouter's documented route limit differs from the 1-million-token model specification, so the latter should not be assumed to apply to every serving route.[4]

Mistral reports training on 3,800 NVIDIA Grace Blackwell GPUs in its European data centers, which serve the preview, and training data spanning more than 160 languages.[2]

## API access and pricing

Mistral's Chat Completions example uses `mistral-large-4` with `reasoning_effort="high"`.[3]

[OpenRouter](https://aiwiki.ai/wiki/openrouter) identifies its route as `mistralai/mistral-large-4-0`.[4] These are provider-specific identifiers, not interchangeable request strings.

The official pricing page displayed the following standard API rates on October 7, 2026, in US dollars per million tokens.[5] Mistral's changelog describes the launch offer as a 50% discount for two weeks.[1]

| Token type | Listed standard rate | Launch sale rate |
|---|---:|---:|
| Input | $1.36 | $0.68[5] |
| Cached input | $0.14 | $0.07[5] |
| Output | $4.18 | $2.09[5] |

These figures describe the launch pricing snapshot, not a permanent discounted rate. Mistral's pricing page separates standard, batch, and priority tiers.[5]

## Tool use and document workflows

The card lists chat completions, structured outputs, function calling, Document QnA, prefix completion, batching, Agents and Conversations, and built-in tools.[3]

[Function calling](https://aiwiki.ai/wiki/function_calling) lets an application provide tool definitions and receive a selected function with generated arguments. In the ordinary Chat Completions workflow, the application executes that function and supplies the result back to the model. Mistral separately offers server-side tools through its Agents and Conversations APIs. Tool-calling support does not mean that an ordinary completion has unrestricted access to an application's systems.[7]

For structured output, Mistral provides JSON mode and custom schemas. Its documentation recommends custom structured outputs when a specific format is required; JSON mode still requires the prompt to request JSON and describe the intended format.[8]

Document QnA combines an OCR stage, which extracts document text and structure, with a language model that answers questions about the extracted content. The documented workflow covers information extraction, summarization, and comparisons across documents.[9]

Agents store a model selection, instructions, tools, and completion settings. Conversations record interactions and can be started with either an agent or a model directly. Mistral's built-in tool types include web search, a code interpreter, image generation, and a document library.[10] In particular, a tool for image generation is a platform capability, not evidence that the model's own output modality includes generated images.

## Evaluations

### Launch results reported by Mistral

The launch announcement reports these preview results.[2]

| Evaluation | Reported result |
|---|---:|
| DeepSWE v1.1 | 61.7%[2] |
| Terminal-Bench 4 | 28.3%[2] |
| Cybench | 93%[2] |
| Lakera B3 attack resistance | 93.3%[2] |

The original [Cybench](https://aiwiki.ai/wiki/cybench) research describes 40 capture-the-flag tasks from four competitions, with command-execution environments and optional subtask guidance. It also studies different agent scaffolds.[15] Its original experiments predate Mistral Large 4; the paper defines the benchmark rather than independently validating this model's launch score.

### Independent evaluator snapshots

Artificial Analysis listed Mistral Large 4 Preview with an Intelligence Index score of 38 on October 7, 2026. Its page identifies the index version as 4.3.2, combining ten evaluations, and reports 116.1 output tokens per second for Mistral's API. It labels the preview proprietary because the weights are not publicly available, in contrast to Mistral's description of the planned release as open-weight.[11]

Vals AI separately published the following results for Mistral Large 4. The displayed uncertainty values are retained with the scores.[12]

| Vals evaluation | Score |
|---|---:|
| Vals Index | 48.05% +/- 1.11[12] |
| Finance Agent v2 | 54.68% +/- 0.58[12] |
| Harvey's Legal Agent Benchmark | 15.83% +/- 2.96[12] |
| Terminal-Bench 4.0 | 22.73% +/- 0.88[12] |

Vals' methodology uses controlled evaluation harnesses, private test sets for proprietary benchmarks, and published standard errors. For benchmarks with multiple runs, including Terminal-Bench and Finance Agent v2, it calculates uncertainty over run-level scores. Such error bars do not capture every change in prompts or deployment settings.[13] Its Terminal-Bench result is separate from Mistral's reported 28.3%; the two figures should not be collapsed into one score.

Harvey's Legal Agent Benchmark evaluates legal deliverables using file-system tools and skills for documents, presentations, and spreadsheets. Vals reports results on a held-out task set and distinguishes whole-task success from satisfying individual criteria.[14] A task pass rate on this benchmark is not a general measure of legal accuracy or an authorization to rely on generated legal advice.

## Safety evaluation scope

Mistral's launch report discusses resistance to indirect [prompt injection](https://aiwiki.ai/wiki/prompt_injection) and refusals of malicious cybersecurity requests, including tests drawn from JailbreakBench, StrongREJECT, and AgentHarm.[2] These evaluations address specified attack or misuse settings.

[AgentHarm](https://aiwiki.ai/wiki/agentharm) contains 110 malicious agent tasks, expanded to 440 through augmentations, across 11 harm categories. Its research evaluates both refusals and whether an attacked agent can complete a harmful multistep workflow.[16] A refusal measurement and an attack-resistance measurement test different behavior; neither constitutes a guarantee that an agent is safe in every deployment.

## References

1. Mistral AI. ["Changelog," October 6, 2026 entry](https://docs.mistral.ai/resources/changelogs). Accessed October 7, 2026.
2. Mistral AI. ["Introducing Mistral Large 4"](https://mistral.ai/news/mistral-large-4/). October 6, 2026. Accessed October 7, 2026.
3. Mistral AI. ["Mistral Large 4" model card, including Weights and Usage tabs](https://docs.mistral.ai/models/mistral-large-4-0). October 6, 2026. Accessed October 7, 2026.
4. OpenRouter. ["Mistral: Mistral Large 4"](https://openrouter.ai/mistralai/mistral-large-4-0). Accessed October 7, 2026.
5. Mistral AI. ["Pricing"](https://docs.mistral.ai/inference/pricing). Accessed October 7, 2026.
6. Mistral AI. ["Model lifecycle policy"](https://docs.mistral.ai/inference/model-lifecycle). Accessed October 7, 2026.
7. Mistral AI. ["Function Calling"](https://docs.mistral.ai/studio/conversations/function-calling). Accessed October 7, 2026.
8. Mistral AI. ["Structured Outputs"](https://docs.mistral.ai/studio/conversations/structured-output). Accessed October 7, 2026.
9. Mistral AI. ["Document AI QnA"](https://docs.mistral.ai/studio/document-processing/document_qna). Accessed October 7, 2026.
10. Mistral AI. ["Agents & Conversations"](https://docs.mistral.ai/studio/agents/agents-api). Accessed October 7, 2026.
11. Artificial Analysis. ["Mistral Large 4 Preview: Intelligence, Performance & Price Analysis"](https://artificialanalysis.ai/models/mistral-large-4). Accessed October 7, 2026.
12. Vals AI. ["Mistral Large 4 Benchmarks, Cost and Capabilities"](https://www.vals.ai/models/mistralai_mistral-large-4). Accessed October 7, 2026.
13. Vals AI. ["Methodology"](https://www.vals.ai/methodology). Accessed October 7, 2026.
14. Vals AI. ["Harvey's Legal Agent Benchmark"](https://www.vals.ai/benchmarks/hlab). Updated October 6, 2026. Accessed October 7, 2026.
15. Zhang, Andy K., et al. ["Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models"](https://arxiv.org/abs/2408.08926). arXiv:2408.08926, submitted August 15, 2024; revised April 12, 2025. ICLR 2025.
16. Andriushchenko, Maksym, et al. ["AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents"](https://arxiv.org/abs/2410.09024). arXiv:2410.09024, submitted October 11, 2024; revised April 18, 2025. ICLR 2025.

