Mistral Large 2 (2407)
Mistral AI's flagship 123B model specialized in multilingual reasoning, precision code generation, and agentic tool use.
128k
Tokens
8.192k
Output limit
Dense Transformer
Model family
123B
Total / Active
$2.00
Per 1M tokens
$6.00
Per 1M tokens
Model Overview
Mistral Large 2 (123B parameters) delivers frontier-level performance across 80+ programming languages and dozens of natural languages with 128k context window and strict adherence to system instructions.
Developer Implementation Notes
Weights available for research and non-commercial use on Hugging Face, commercial use via Mistral La Plateforme and cloud partners.
Key Strengths
- State-of-the-art multilingual reasoning in French, German, Spanish, and English
- Superior agentic tool-calling precision and JSON schema output
- Large 128k context window
- Very strong code generation across 80+ languages
Limitations & Boundaries
- 123B size is heavy for dual-GPU consumer setups
- Commercial self-hosting requires commercial agreement
Capabilities & Modalities
Best Production Use Cases
- European enterprise compliance and multilingual deployment
- Agentic workflow orchestration with multiple tool calls
- Full-stack software engineering
Verified Benchmark Results
Standardized evaluations with methodology notes and authoritative citation links.
Pricing & Inference Cost Calculator
Token Cost Estimator – Mistral Large 2 (2407)
Calculate projected inference spend with prompt caching
Quick Workload Presets
$4.0000
2.00M input tokens
$3.0000
0.50M output tokens
$7.00
Avg: $0.00700 / req
Compare with Similar Models
DeepSeek-R1
Open-weights frontier reasoning model trained via large-scale reinforcement learning without supervised cold start.
Codestral 2501
Mistral AI's dedicated code generation model with 256k context window and fill-in-the-middle (FIM) capabilities.
DeepSeek-V3
Frontier open-weights 671B MoE base and chat model with multi-head latent attention (MLA) and dual-pipe training.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.