The call
Mistral released Large 4 as a public API preview on 6 October and promised open weights by the end of the month (TechCrunch). It reached 1,591 points and 964 comments on Hacker News within a day (independent, HN digest). A second open-weight launch, Reflection’s 501B-parameter Beam, landed the day before.
What it is, and is not
It is a 1.05 trillion-parameter mixture-of-experts model that uses 49 billion parameters per token, about 4.7% of its weights (vendor, via MarkTechPost). It reads text and images. Artificial Analysis scores it 38 on its Intelligence Index, level with OpenAI’s GPT-6 Luna and one point behind DeepSeek V4.1 Flash (independent, Artificial Analysis). The preview API lists $1.36 per million input tokens and $4.18 per million output, half price for the first two weeks (independent, same source).
It is not, yet, a private model. Today every token goes to Mistral’s servers. The weights are dated 31 October on the Hugging Face card and 27 October in The Next Web. Before release, Mistral is red-teaming a less moderated version “with cybersecurity leaders, vetted partners and state authorities” (vendor, via TestingCatalog). Open-weight is a promise with a date, not a deployment option.
The number nobody is quoting
Coverage quotes the 4.7%: only 49 billion parameters work on each token, so it runs cheaply. That is compute. Memory is different. To serve the model you hold all 1.05 trillion parameters in GPU memory, every hour, whether a token uses them or not.
| Figure | Value | Label |
|---|---|---|
| Weights at 8-bit (1 byte per parameter) | 1.05T × 1 = about 1,050 GB | our maths |
| Weights at 4-bit (0.5 bytes per parameter) | 1.05T × 0.5 = about 525 GB | our maths |
| Mistral’s serving guidance | 4 to 8 Nvidia B200 or B300 GPUs | vendor, via AlphaSignal |
| Four B200 (180 GB each) | 720 GB: 4-bit fits with 195 GB spare; 8-bit is 330 GB short | memory independent; fit is our maths |
| Four B300 (268 GB each) | 1,072 GB: 8-bit fits with 22 GB spare, about 2% | memory independent; fit is our maths |
| Eight H100 (80 GB each), the common enterprise node | 640 GB: 8-bit is 410 GB short, so two nodes | memory vendor; fit is our maths |
| Reflection Beam, 501B parameters, at 8-bit | about 501 GB: fits one eight-H100 node with 139 GB spare | parameters vendor; fit is our maths |
“Four GPUs” is true at 4-bit on B200s, or at 8-bit on B300s with almost nothing left for long prompts. These are weights only; working memory for prompts comes on top. Check which generation you own before anyone promises a date.
Second number. At the same Intelligence Index score, Luna lists $0.10 input and $0.50 output (OpenAI). Mistral is 13.6x Luna on input and 8.4x on output at list, and 6.8x and 4.2x during the launch discount (our maths). You pay that premium for one thing: the option to run it yourself. So the licence is the product.
Missing:
- The licence. Commercial use, modification, and redistribution terms are unpublished (AlphaSignal). Beam, by contrast, is Apache 2.0 (Reflection).
- A serving specification: precision, memory, throughput per node, and context length when self-hosted.
- One release date. 27 or 31 October?
- One context length. 524,288 tokens on OpenRouter; 1 million in MarkTechPost.
- Data retention and region terms for the preview API.
- An SLA. OpenRouter shows 96.48% uptime for the single Mistral endpoint (independent).
- Independent task results beyond one index. Benchmarks so far are vendor tables.
Who owns what
Weights are Mistral’s job. Hardware, serving, and output quality are ours.
In a Salesforce estate, a self-hosted model plugs in through the LLM Open Connector in Einstein Studio Model Builder, which accepts any model behind an endpoint that follows its specification and exposes it to Prompt Builder and the Models API. Put the endpoint behind a Named Credential, ground it with Data Cloud, and log every call. The model is a component. The evaluation set, the fallback, and the audit trail are what make it private in practice.
Monday morning
- Step 1. List the workloads that cannot leave your estate, by data class. If the list is empty, the hosted API is cheaper elsewhere. Owner: CIO with the CISO.
- Step 2. Count the GPU memory you own or can reserve, by GPU type. Set it against 525 GB and 1,050 GB. Owner: infrastructure lead.
- Step 3. Book the licence review for release day. No licence, no production. Owner: general counsel.
- Step 4. Wire the LLM Open Connector in a sandbox with an open model you already run, so a swap is configuration, not a project. Owner: Salesforce platform owner.
- Step 5. If you test the preview API, use public or synthetic data only. Owner: AI lead.
What would change our call: we move this to PILOT when the weights ship with a licence that permits commercial self-hosting and Mistral publishes a serving specification.
Living page. Last updated 7 Oct 2026. What would change our call: weights released under a commercial licence, with a published serving spec.
Frequently asked questions
Can we run Mistral Large 4 on our own servers today?
No. Today it is a public preview on Mistral’s hosted API. Mistral says the weights follow at the end of October, after safety testing, and has not yet published the licence or a serving specification.
How much GPU memory does Mistral Large 4 need?
Our arithmetic: about 1,050 GB for the weights alone at 8-bit precision and about 525 GB at 4-bit, before any working memory for long prompts. Mistral has said four to eight Nvidia B200 or B300 GPUs. Four B200s hold the 4-bit weights; they do not hold the 8-bit weights.
Can Salesforce use a self-hosted open-weight model?
Yes, for generative features. Salesforce’s LLM Open Connector in Einstein Studio Model Builder connects any model behind an endpoint that follows its specification, and the model then appears in Prompt Builder and the Models API. Keep the endpoint behind a Named Credential.