Mistral Large 4 Weights, Download & Deployment Guide
Try Mistral Large 4 through the online Playground and API, and understand weights, licensing and hardware requirements before planning a local deployment.
Last updated: 2026-10-07
Can I download Mistral Large 4 now?
As checked on October 7, 2026, Mistral Large 4 is available in public API preview. The weights are planned for release by the end of October. The preview API is available before the planned weights release; those are separate milestones.
You can already try Mistral Large 4 in the Playground without downloading weights or buying hardware. Sign in to use the playground for free. The planned month-end weights release does not guarantee a particular date.
How can I use it while waiting for the weights?
- Online playground: sign in and try a prompt without setting up a server. Use the playground for free, then choose a credit pack or monthly plan for continued use.
- API integration: create an account key and follow the API guide. The public model ID is
mistral-large-4-0. Use it with the base URL and account key.
A key from one service does not authenticate you to another. Hosted access also does not install the model on your computer.
Does “open weights” mean open source or free commercial use?
Open weights describe access to trained model parameters. The license determines permissions and conditions for use, modification and redistribution. Availability of weights alone does not establish that all training code and data are open, or that every commercial use is permitted without conditions.
For Large 4, consult the license attached to its actual release before making a commercial deployment decision. Do not inherit a license from Large 3 or another model. If a requirement is unclear, review the released terms with your legal team.
Downloading weights also does not make compute free. A deployment still needs hardware or cloud capacity, storage, electricity and ongoing operation.
Can I run it on my laptop or one GPU?
Do not plan on a normal laptop or single consumer GPU running the full model. The official materials describe a model at roughly trillion-parameter scale. Its mixture-of-experts architecture activates only part of the model for a token, but that active count is not the total weight-storage requirement.
We do not yet have a verified minimum GPU configuration for the planned release. A useful estimate needs the released weight format, precision, supported runtime and target workload. Quantization may reduce storage requirements, but its availability and quality must be verified for the exact release.
Measure more than whether the weights load: long-context cache, concurrent requests, image processing and runtime overhead all affect capacity. A deployment that fits one short request may not meet your intended throughput or context length.
Will Ollama, vLLM or other runtimes support it?
Do not assume compatibility from the model name. After the weights are released, check the release-specific deployment instructions and the runtime’s support for the exact architecture, tokenizer, vision inputs and weight format.
A useful deployment guide should specify a repository revision, runtime version, precision, hardware and measured limits. Until those are confirmed, a generic installation command can be misleading. This page intentionally separates planning guidance from tested deployment instructions.
Hosted API or self-hosting: which should I choose?
Start with the hosted Playground or API when you want to test quality without operating the model. The documentation explains supported inputs and request limits. Use your own tasks to compare quality, response time and total cost, including retries.
Evaluate self-hosting after the release if you need control over infrastructure or operation and can support the required hardware. Compare sustained utilization, administration and reliability costs against hosted usage. Low-volume usage does not automatically become cheaper with self-hosting.
The pricing explanation covers free playground access, credit usage, one-time packs and monthly plans. Use the model-selection guide to plan an evaluation with your own tasks.
Does self-hosting solve every privacy requirement?
Self-hosting can put inference infrastructure under your control, but the complete application still matters: document storage, external tools, telemetry, logs and access permissions can all move or retain data.
Playground requests are processed by a hosted model provider. Infrastructure, processing regions and contract terms depend on the hosting configuration. Read the privacy policy and confirm retention, training use, processing region and required agreements with support before sending confidential company material.
What should I verify when the weights arrive?
- Follow a repository link from the official model card or announcement.
- Read the release-specific license and model instructions.
- Confirm runtime support and budget for weights, runtime overhead and your intended context length.
- Test text, images, structured output and tool calls that your application actually needs.
- Measure latency, throughput, quality and operating cost before replacing an existing service.
For hosted integration now, start with the API and playground documentation. For integration help, contact support.