metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
Mistral
Open Weights
Mixture of Experts
Sovereign AI
Nvidia

Mistral Large 4: 1 Trillion Parameters, Open Weights Due

Mistral Large 4 preview is a 1T-parameter MoE with 49B active. What that means for self-hosting cost, rival open models, EU sovereignty and licence terms.

Metir AI TeamOctober 7, 20268 min read
Mistral Large 4: 1 Trillion Parameters, Open Weights Due

On October 6, 2026, Mistral announced a public preview of Mistral Large 4, which it calls its largest and most capable model to date. The company describes a mixture-of-experts system with 1 trillion total parameters and 49 billion active per token, natively multimodal, and trained on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European datacenters. The weights are promised by the end of October, which means that today Mistral Large 4 is reachable only through Mistral's API. This article separates what is confirmed from what is still reported, then explains what a 1T-total, 49B-active design means for anyone considering self-hosting.

Mistral AI logoMistral AI
NVIDIA logoNVIDIA
DeepSeek logoDeepSeek
Moonshot AI logoMoonshot AI
Qwen logoQwen
Mistral Large 4 and the open-weight models it will be compared against.
1TTotal parametersMistral announcement
49BActive per tokenAbout 5% of weights
160+LanguagesIncludes every official EU language
$1.36 / $4.18Per 1M input / output tokensPreview API pricing

What Mistral Large 4 is, and what is still unconfirmed

Mistral's announcement confirms the headline specifications: 1 trillion total parameters, 49 billion active, native multimodality, support for more than 160 languages including every official language of the European Union, and preview access through Mistral Studio at $1.36 per million input tokens and $4.18 per million output tokens. It also says the model was trained from scratch and that deployment will be available in several regions worldwide, including a European deployment operated end to end by Mistral.

Several details differ between sources, and a few circulating claims are not in any source we could read:

  • GPU count. Mistral's page says 3,800 Grace Blackwell GPUs. The Next Web reports roughly 4,000 GPUs, about two months of training and around 10 megawatts of power. The figures are consistent with each other as rounding, but 3,800 is Mistral's own number.
  • Release date. Mistral says "end of October." The Next Web gives October 27, 2026; Implicator also cites October 27. A date of October 31 that appeared in some early chatter is not supported by what we read.
  • Context window. A roughly 1 million token context has been mentioned informally, but neither Mistral's announcement nor the coverage we read states a context length. We treat it as unconfirmed.
  • Licence. Mistral's page does not state the licence. Implicator expects a custom Mistral licence, in contrast to the Apache 2.0 licence on the previous model, but that is an expectation, not an announcement.
“

Mistral calls ML4 open-weight, not open source in the full software sense. Wait for the actual weight release and license before assuming commercial or redistribution rights.

Developers Digest, Oct 6, 2026

Active versus total parameters: what the 49B actually buys

In a mixture-of-experts (MoE) model, the network is split into many specialist sub-networks, and a router sends each token to only a few of them. Total parameters measure everything the model stores. Active parameters measure what any single token actually computes through. For Mistral Large 4, 49 billion of 1 trillion is about 4.9% of the weights per token.

That ratio drives two different costs:

  • Compute per token scales with active parameters. A dense 1T model would do roughly 20 times the arithmetic per token (1,000 divided by 49, our own calculation). Mistral Large 4 does the work of a model near 49B per token.
  • Memory scales with total parameters. Every expert must be reachable, because the next token may route to any of them. The memory bill is that of a 1T model.

Total vs active parameters in open-weight MoE models

Billions of parameters, as reported by each lab or the cited coverage. Active parameters set compute per token; total parameters set memory.

Sources: Mistral (Oct 6, 2026), Implicator, Reflection AI, MarkTechPost, MorphLLM. Kimi K3 and Qwen 3.8 Max omitted: active counts not sourced. Hover a bar for details.

This is the reason sparse models are cheap to call through an API and demanding to host. The provider amortises one very large memory footprint across thousands of concurrent requests. A single organisation running its own copy pays for that footprint alone.

Self-hosting a 1T model: the hardware arithmetic

Weight storage is parameters multiplied by bytes per parameter. At 16-bit precision, 1 trillion parameters is about 2 TB. At 8-bit it is about 1 TB, and at 4-bit about 500 GB. These are our own calculations from the reported parameter count, and they exclude the key-value cache that grows with context length and concurrent users, plus runtime overhead. Mistral has not said which precision the weights will ship in.

Approximate memory for a 1T-parameter model's weights

Gigabytes, our own arithmetic from the reported 1 trillion parameters. Excludes KV cache and runtime overhead. The shipped precision has not been announced.

Source: Mistral (1T total, 49B active); byte sizes are standard numeric formats; GB figures are our arithmetic. Hover a bar for details.

Even the most aggressive case exceeds any laptop or single workstation GPU, which matches Implicator's note that the model cannot run on a desktop or laptop and may be impractical for some universities. Realistically, self-hosting means a multi-GPU server or several, with fast interconnect so that experts spread across devices can exchange activations. Developers Digest makes the same point from the operations side: a 1T sparse model with 49B active parameters is not the same operational object as a smaller model, and serving topology, memory bandwidth, quantization quality and batching all need evaluation.

Two engineering consequences follow. First, decoding is usually limited by memory bandwidth rather than arithmetic, so the small active slice helps less than the headline suggests at low batch sizes. Second, at high batch sizes different tokens in a batch hit different experts, so most of the model is touched on every step and the compute advantage returns. In practice the model favours shared, busy deployments over a single-user box.

How Large 4 compares with other open-weight MoE models

Raw size is the easiest comparison and the least informative. The list below uses parameter counts from our earlier coverage of each model.

  • DeepSeek V4 Pro: about 1.6 trillion total and about 49 billion active, MIT licence. See our DeepSeek V4 analysis. It shares Mistral's active count with a larger total, so it is sparser.
  • Kimi K3: 2.8 trillion total, described by Moonshot as the first open model in the 3-trillion class. See our Kimi K3 analysis. We did not find a sourced active-parameter figure.
  • Qwen 3.8 Max: 2.4 trillion total, multimodal, per our Qwen coverage, also without a sourced active count.
  • Reflection Beam: 501 billion total and 23 billion active, per our Beam analysis.
  • StepFun Step 5 Preview: 600 billion total and 27 billion active, with open weights promised.
  • Mistral Large 3: 675 billion total and 41 billion active, Apache 2.0, according to Implicator.

Benchmark comparisons are harder. Mistral reports state-of-the-art open-weight results on enterprise workloads such as cybersecurity, finance and manufacturing, and Implicator lists 61.7% on DeepSWE v1.1 and 93% on Cybench. Coverage disagrees on rivals: The Next Web puts GLM-5.3 at 61% and DeepSeek-V4-Pro at 57% on DeepSWE, while Implicator says GLM-5.3 and Kimi K3 sit near 69% on the DeepSWE leaderboard, and roughly 74% for closed models. These cannot both be right on the same configuration, so we do not draw a ranking. Developers Digest likewise calls the figures claims to verify, not a settled ranking. Independent evaluation after the weights land will be more informative.

The European sovereignty angle

Mistral's pitch is that open weights let an organisation keep control of the model it depends on. Chief scientist Guillaume Lample argued that you cannot afford to be vulnerable to the model you use to protect yourself disappearing one day, or being too limited. CEO Arthur Mensch said the narrative that Europe cannot compete is not true. The Next Web ties the launch to Mistral's September €3 billion Series D and 125-plus enterprise customers, including Airbus, ASML and HSBC. For background, see our coverage of the Samsung-led round.

Arthur Mensch, co-founder and CEO of Mistral AI, speaking on stage at Techarena 2026
Arthur Mensch, Mistral AI co-founder and CEO, speaking at Techarena 2026 (photo taken February 11, 2026, months before the Large 4 announcement). Photo: Jan2342342423, CC0, via Wikimedia Commons.

Sovereignty here has three separable layers. Training location (European datacenters, per Mistral) is one. Deployment (an end-to-end European service operated by Mistral) is another. Weight custody, meaning the ability to run the model on your own infrastructure under your own legal regime, is the third, and it depends entirely on the licence and the promised release. Implicator makes the dependency explicit: until the weights ship, using Large 4 means relying on Mistral's API, the same vendor dependence a closed model imposes.

Licence terms: the open question

"Open weights" describes downloadable parameters, not a legal permission set. Apache 2.0 permits commercial use, modification and redistribution. A custom licence could add thresholds, acceptable-use rules or limits on competing services; DeepSeek V4's MIT licence sits at the permissive end. Mistral has not published Large 4's terms, so the practical advice from multiple outlets is the same: read the licence on release day before building redistribution or resale plans. Our earlier piece on Mistral's open-weight MoE strategy explains why licence can matter more than the leaderboard.

What to watch before the open-weight release

  • Whether weights appear on schedule and in which precision (BF16, FP8 or lower).
  • The licence text, especially commercial and redistribution terms.
  • A stated context length, and independent long-context tests.
  • Third-party evaluations of the DeepSWE, Cybench and Harvey figures.
  • Serving recipes from inference engines, which determine real cost per token off Mistral's own API.

For teams comparing models like these, a model-agnostic workspace such as Metir makes it easier to test a newly released open model against existing choices without rebuilding a workflow.

Sources:

  • Introducing Mistral Large 4 | Mistral AI
  • Europe's Mistral launches Large 4 to challenge China's lead in open AI models | The Next Web
  • Mistral Large 4 Preview Pitches Open Weights Against Lock-In | Implicator
  • Mistral Large 4 Is the Open-Weight Frontier Test | Developers Digest
  • Mistral Large 4 Preview: 1 Trillion Parameters, Open Weights at End of October | iMasters

Image credits

  • Hero: "Prime Minister Keir Starmer meets Arthur Mensch," 10 Downing Street, London, January 9, 2025. Photo: Simon Dawson / No 10 Downing Street, via Wikimedia Commons, licensed under the Open Government Licence v3.0. The photo predates and is unrelated to the Large 4 launch.
  • In-article: "Arthur Mensch" at Techarena 2026, February 11, 2026. Photo: Jan2342342423, via Wikimedia Commons, CC0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational