Mistral Large 4
Mistral Large 4 is a general-purpose multimodal model developed by Mistral AI. It entered public preview on October 6, 2026, with a mixture-of-experts architecture and API access. The initial release preceded the planned distribution of its weights.[1] It belongs to the Mistral Large family, but its preview specifications and availability should not be confused with those of Mistral Large 3.
Release and availability
Mistral called the model ML4 and "le Chonk", offered a Studio preview API, and announced weights for the end of October 2026.[2]
On October 7, the model card listed weights and their license as forthcoming, not an available checkpoint with published licensing terms.[3]
Public preview is a distinct stage in Mistral's model lifecycle. Mistral permits silent updates during this stage and does not guarantee that a preview will reach general availability. Results obtained from the preview can therefore describe a changing service rather than a permanently fixed checkpoint.[6]
Architecture and specifications
The specifications below describe the October 2026 preview.[3]
| Property | Published value |
|---|---|
| Architecture | Granular mixture of experts[3] |
| Total parameters | 1.05 trillion in the model card[3] |
| Vision encoder | 1.6 billion parameters[3] |
| Active parameters | 52 billion in the English model card; 49 billion in the launch announcement[3][2] |
| Mistral context window | 1 million tokens in the model card[3] |
| OpenRouter context limit | 512K tokens, displayed as 524K in its summary[4] |
| OpenRouter maximum output | 256K tokens[4] |
| Input and output | Text and image input; text output on OpenRouter[4] |
The English card lists 52B active parameters; the announcement states 49B. These are conflicting published figures, not a reconciled count.[3][2]
OpenRouter's documented route limit differs from the 1-million-token model specification, so the latter should not be assumed to apply to every serving route.[4]
Mistral reports training on 3,800 NVIDIA Grace Blackwell GPUs in its European data centers, which serve the preview, and training data spanning more than 160 languages.[2]
API access and pricing
Mistral's Chat Completions example uses mistral-large-4 with reasoning_effort="high".[3]
OpenRouter identifies its route as mistralai/mistral-large-4-0.[4] These are provider-specific identifiers, not interchangeable request strings.
The official pricing page displayed the following standard API rates on October 7, 2026, in US dollars per million tokens.[5] Mistral's changelog describes the launch offer as a 50% discount for two weeks.[1]
These figures describe the launch pricing snapshot, not a permanent discounted rate. Mistral's pricing page separates standard, batch, and priority tiers.[5]
Tool use and document workflows
The card lists chat completions, structured outputs, function calling, Document QnA, prefix completion, batching, Agents and Conversations, and built-in tools.[3]
Function calling lets an application provide tool definitions and receive a selected function with generated arguments. In the ordinary Chat Completions workflow, the application executes that function and supplies the result back to the model. Mistral separately offers server-side tools through its Agents and Conversations APIs. Tool-calling support does not mean that an ordinary completion has unrestricted access to an application's systems.[7]
For structured output, Mistral provides JSON mode and custom schemas. Its documentation recommends custom structured outputs when a specific format is required; JSON mode still requires the prompt to request JSON and describe the intended format.[8]
Document QnA combines an OCR stage, which extracts document text and structure, with a language model that answers questions about the extracted content. The documented workflow covers information extraction, summarization, and comparisons across documents.[9]
Agents store a model selection, instructions, tools, and completion settings. Conversations record interactions and can be started with either an agent or a model directly. Mistral's built-in tool types include web search, a code interpreter, image generation, and a document library.[10] In particular, a tool for image generation is a platform capability, not evidence that the model's own output modality includes generated images.
Evaluations
Launch results reported by Mistral
The launch announcement reports these preview results.[2]
The original Cybench research describes 40 capture-the-flag tasks from four competitions, with command-execution environments and optional subtask guidance. It also studies different agent scaffolds.[15] Its original experiments predate Mistral Large 4; the paper defines the benchmark rather than independently validating this model's launch score.
Independent evaluator snapshots
Artificial Analysis listed Mistral Large 4 Preview with an Intelligence Index score of 38 on October 7, 2026. Its page identifies the index version as 4.3.2, combining ten evaluations, and reports 116.1 output tokens per second for Mistral's API. It labels the preview proprietary because the weights are not publicly available, in contrast to Mistral's description of the planned release as open-weight.[11]
Vals AI separately published the following results for Mistral Large 4. The displayed uncertainty values are retained with the scores.[12]
Vals' methodology uses controlled evaluation harnesses, private test sets for proprietary benchmarks, and published standard errors. For benchmarks with multiple runs, including Terminal-Bench and Finance Agent v2, it calculates uncertainty over run-level scores. Such error bars do not capture every change in prompts or deployment settings.[13] Its Terminal-Bench result is separate from Mistral's reported 28.3%; the two figures should not be collapsed into one score.
Harvey's Legal Agent Benchmark evaluates legal deliverables using file-system tools and skills for documents, presentations, and spreadsheets. Vals reports results on a held-out task set and distinguishes whole-task success from satisfying individual criteria.[14] A task pass rate on this benchmark is not a general measure of legal accuracy or an authorization to rely on generated legal advice.
Safety evaluation scope
Mistral's launch report discusses resistance to indirect prompt injection and refusals of malicious cybersecurity requests, including tests drawn from JailbreakBench, StrongREJECT, and AgentHarm.[2] These evaluations address specified attack or misuse settings.
AgentHarm contains 110 malicious agent tasks, expanded to 440 through augmentations, across 11 harm categories. Its research evaluates both refusals and whether an attacked agent can complete a harmful multistep workflow.[16] A refusal measurement and an attack-resistance measurement test different behavior; neither constitutes a guarantee that an agent is safe in every deployment.
References
- ^1 ^2Mistral AI. "Changelog," October 6, 2026 entry. Accessed October 7, 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Mistral AI. "Introducing Mistral Large 4". October 6, 2026. Accessed October 7, 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Mistral AI. "Mistral Large 4" model card, including Weights and Usage tabs. October 6, 2026. Accessed October 7, 2026.
- ^1 ^2 ^3 ^4 ^5OpenRouter. "Mistral: Mistral Large 4". Accessed October 7, 2026.
- ^1 ^2 ^3 ^4 ^5Mistral AI. "Pricing". Accessed October 7, 2026.
- ^Mistral AI. "Model lifecycle policy". Accessed October 7, 2026.
- ^Mistral AI. "Function Calling". Accessed October 7, 2026.
- ^Mistral AI. "Structured Outputs". Accessed October 7, 2026.
- ^Mistral AI. "Document AI QnA". Accessed October 7, 2026.
- ^Mistral AI. "Agents & Conversations". Accessed October 7, 2026.
- ^Artificial Analysis. "Mistral Large 4 Preview: Intelligence, Performance & Price Analysis". Accessed October 7, 2026.
- ^1 ^2 ^3 ^4 ^5Vals AI. "Mistral Large 4 Benchmarks, Cost and Capabilities". Accessed October 7, 2026.
- ^Vals AI. "Methodology". Accessed October 7, 2026.
- ^Vals AI. "Harvey's Legal Agent Benchmark". Updated October 6, 2026. Accessed October 7, 2026.
- ^Zhang, Andy K., et al. "Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models". arXiv:2408.08926, submitted August 15, 2024; revised April 12, 2025. ICLR 2025.
- ^Andriushchenko, Maksym, et al. "AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents". arXiv:2410.09024, submitted October 11, 2024; revised April 18, 2025. ICLR 2025.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 1,321 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent full-article review against 16 primary and research references, October 7, 2026. Checked subject identity, specifications, availability, limitations and citation support.
Cite this page: AI Wiki. "Mistral Large 4." aiwiki.ai, updated 6 Oct 2026, fact-checked 6 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/mistral_large_4