Docs

Mistral Large 4 API & Playground Guide

Get started with the playground and API. Find model IDs, request examples, service limits, free allowances and troubleshooting steps.

Open the playground

Overview

Use the playground after signing in, or connect your application with an account API key. The endpoint uses model ID mistral-large-4-0. Create your key in account settings and use the base URL and examples below. The published model specification lists a 1M-token context window. The current integration is configured around 524,288 context tokens and accepts max_tokens up to 262,144, subject to upstream support and the remaining context. The playground has smaller limits described below. Supported inputs are text and image URLs, not direct PDF, video or audio uploads.

Your first conversation

The playground is free to use after sign-in. Choose an example or write your own prompt to get started. Responses within the free allowance stream progressively; additional requests use credits and return answers after credit settlement.

  1. Sign in to your account.

  2. Describe a task and its context.

  3. Read, follow up and refine.

Write a useful prompt

State the task, supply the relevant source and define the output. For writing, name the audience, language, tone and terminology. For document questions, label passages and ask for quotations from those passages; require the model to say when evidence is missing. Extract PDF text before pasting it here. For code, include the error and expected behavior. Verify claims and run code before relying on the result.

Goal: explain this function to a new teammate.
Context: a small TypeScript application.
Output: a short explanation, potential edge cases, and one improvement.

Responses and history

Press Enter to send, or Shift + Enter for a new line. Stop interrupts the current response. Clear conversation resets this playground. Conversation context is kept while the page is open; refreshing the page starts a fresh conversation. Copy any answers you want to keep before leaving.

Model API

Create an API key in Account → API Keys. Install the Python client with pip install openai and store your account key in the MODEL_API_KEY server environment variable. Use the base URL and model ID below. POST /api/v1/chat/completions requires Authorization: Bearer <your-account-api-key> and Content-Type: application/json. Non-streaming responses contain choices[0].message and usage. The SDK example runs on your server; keep the key out of browser code.

POST https://mistrallarge4.com/api/v1/chat/completions

model: mistral-large-4-0

Create an account API key
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://mistrallarge4.com/api/v1",
    api_key=os.environ["MODEL_API_KEY"],
)

response = client.chat.completions.create(
    model="mistral-large-4-0",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=2048,
    extra_body={"reasoning": {"effort": "none"}},
)
print(response.choices[0].message.content)

Request parameters and limits

The API accepts up to 128 messages and a 2 MiB JSON request body. The API and playground share a limit of 10 requests per minute per account on the current server instance. Requests time out after 120 seconds. Output limits include reasoning tokens; raise max_tokens or choose none if reasoning leaves no visible answer. Only the parameters listed here are accepted.

ParameterAccepted values
modelmistral-large-4-0; used by default when omitted.
messagesRequired array of system, user, assistant or tool messages. User content accepts text or text/image_url parts. Tool replies require tool_call_id.
reasoning{ "effort": "none" | "high" }; default none. reasoning_effort is also accepted; send only one of these fields.
temperature0–2. Omit to use the model default. Playground defaults to 0.7.
max_tokens1–262,144; default 2,048. Includes reasoning and visible output, and must fit in the remaining model context.
streamBoolean; default false. With true, optionally send stream_options: undefined.
response_format{ "type": "text" | "json_object" } or undefined }. Ask for JSON in the prompt when using json_object.
toolsUp to 32 function definitions. tool_choice accepts auto, none, required or a named function. Your application executes tools and returns results.
top_p / stop / seed / frequency_penalty / presence_penaltytop_p: 0–1; stop: a string or up to 4 strings; seed: integer; frequency_penalty and presence_penalty: −2 to 2. All are optional.

Streaming responses

Set stream to true to receive SSE data events. Read choices[0].delta.content for text, delta.tool_calls for incremental tool calls and usage for token totals when requested. The last event is data: [DONE]. An interrupted stream can contain an error object; discard incomplete tool arguments. The examples below reuse client from the API example. Raw HTTP clients must buffer chunks until a complete SSE event is available.

stream = client.chat.completions.create(
    model="mistral-large-4-0",
    messages=[{"role": "user", "content": "Explain a retry strategy."}],
    stream=True,
    stream_options={"include_usage": True},
    max_tokens=2048,
)

for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
    if chunk.usage:
        print(chunk.usage)

Images and structured output

Set image_url to an accessible public HTTP or HTTPS image URL without embedded credentials, then combine it with text content. Video and audio content parts are not supported. Use json_object for a JSON object, or json_schema with a named schema for a defined structure. Check finish_reason: length means the output budget was exhausted and the JSON may be incomplete.

response = client.chat.completions.create(
    model="mistral-large-4-0",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this image as JSON."},
            {"type": "image_url", "image_url": {"url": image_url}},
        ],
    }],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "image_summary",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {"summary": {"type": "string"}},
                "required": ["summary"],
                "additionalProperties": False,
            },
        },
    },
)

Tool calling

Define tools and read message.tool_calls. The API returns function names and JSON argument strings; it does not execute your functions. Validate arguments and permissions, execute approved calls in your application, then append the assistant message and a role: tool message with the matching tool_call_id before sending the next request. The assistant in your account is a separate product surface with its own tools and conversation history.

response = client.chat.completions.create(
    model="mistral-large-4-0",
    messages=[{"role": "user", "content": "Check order 123."}],
    tools=[{
        "type": "function",
        "function": {
            "name": "get_order",
            "description": "Look up an order the user is allowed to access.",
            "parameters": {
                "type": "object",
                "properties": {"order_id": {"type": "string"}},
                "required": ["order_id"],
            },
        },
    }],
    tool_choice="auto",
)
print(response.choices[0].message.tool_calls)

Playground session API

POST /api/playground uses your signed-in session and does not require an API key or model ID. GET /api/playground returns the public model ID and supported controls. The playground accepts up to 24 messages, 24,000 characters per message, a 4,000-character system prompt and a 96 KiB request body. maxTokens is 16–16,384; reasoningEffort is none or high; outputFormat is text or json. An image is attached as media: undefined on a user message. This endpoint uses custom SSE events: delta contains content; done contains data.message, data.usage, data.creditsUsed and data.elapsedMs; error contains message. Run the example from the playground page after signing in. Use the API-key endpoint above for external applications.

const response = await fetch('/api/playground', {
  method: 'POST',
  credentials: 'same-origin',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    messages: [{ role: 'user', content: 'Hello!' }],
    systemPrompt: 'Give a clear, concise answer.',
    reasoningEffort: 'none',
    temperature: 0.7,
    maxTokens: 2048,
    outputFormat: 'text',
  }),
});
if (!response.ok) throw new Error((await response.json()).message);
const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  console.log(value); // SSE bytes; events may span multiple chunks.
}

Credits and plans

The playground is free to use after sign-in, subject to a daily allowance that resets at midnight UTC+8 (Asia/Singapore). A validated request counts when it is admitted for model generation, including requests that later fail or are cancelled. Rejected requests do not use the free allowance. After the free allowance is used up, completed answers cost actual model usage in USD × 12,000 credits, rounded up. Paid answers are returned after credit settlement; insufficient balance or missing cost data returns an error without an answer. Failed or cancelled generation does not deduct credits. The server allows one request per user every 3 seconds. API chat and media tools keep their existing policies.

Pricing

Troubleshooting and support

400: check the model ID, messages, parameters or attachments. 401: sign in for the playground or check your account API key for API calls. 402: your daily free playground allowance is used up and your credit balance is insufficient; top up on Pricing or wait for the next day. 413: reduce the request size. 415: send application/json. 429: wait for Retry-After before retrying; the playground allows one request per user every 3 seconds. 502/503: the model service is temporarily unavailable. HTTP errors return undefined. Streaming failures are reported inside the stream. If high reasoning consumes the output budget, increase max_tokens or switch to none. Contact support if the problem continues.

Contact us