Mistral Medium 3.5
The dense 128B parameter powerhouse offering configurable reasoning effort, native vision, and 256k context window.
Generic Info
- Publisher: Mistral AI
- Release Date: 2026
- Architecture: Dense 128B Parameters Transformer with Multimodal Vision Encoder
- Context Window: 256,000 tokens (256k context)
- License: Modified MIT License (Permissive commercial & non-commercial use)
- Key Capabilities: Dual Mode (Instant vs. Reasoning), Image Understanding, First-Class Function Calling, Structured JSON Output
Mistral Medium 3.5 represents Mistral's flagship dense architecture. Built with 128 billion dense parameters, it pairs raw parameter capacity with the agility to switch seamlessly between sub-second conversational replies and deep test-time reasoning. With multimodal input support, a 256k context window, and industry-standard tool calling accuracy, it provides an exceptional backbone for production agent services.
Hello World Guide
Run Mistral Medium 3.5 using vLLM or the official Mistral inference stack.
from vllm import LLM, SamplingParams
model_name = "mistralai/Mistral-Medium-3.5-128B"
# Initialize vLLM with tensor parallelism across GPUs
llm = LLM(
model=model_name,
tensor_parallel_size=4,
max_model_len=32768
)
sampling_params = SamplingParams(
max_tokens=1024,
temperature=0.7
)
prompt = "You are a senior system architect. Design an event-driven distributed pipeline for high-frequency telemetry."
outputs = llm.generate([prompt], sampling_params)
for output in outputs:
print(output.outputs[0].text)
Industry Usage
Dynamic Latency Scaling
Switch between zero-overhead instantaneous chat responses and multi-minute chain-of-thought reasoning depending on task complexity.
Production Agent Pipelines
Best-in-class JSON schema adherence and native tool calling eliminate parsing errors when driving automated backend APIs.
Document & Visual Auditing
256k context combined with vision enables auditing long regulatory PDFs, financial prospectuses, and multi-page schematics.