Qwen 2.5 72B Instruct
Alibaba's flagship open foundation model with world-class multilingual, mathematical, and coding capabilities.
128k
Tokens
8.192k
Output limit
Dense Transformer
Model family
72.7B
Total / Active
$0.00
Per 1M tokens
$0.00
Per 1M tokens
Model Overview
Qwen 2.5 72B Instruct is an open-weights dense transformer trained on 18 trillion tokens with superior mathematical reasoning, 128k context length, and deep multilingual competence across 29+ languages.
Developer Implementation Notes
Apache 2.0 license. Scores 83.1% on MMLU-Pro and 85.5% on MATH-500.
Key Strengths
- Apache 2.0 permissive license
- Exceptional mathematical and reasoning benchmark scores
- Support for 29+ natural languages
- 128k context window
Limitations & Boundaries
- Requires dual 24GB GPUs or single 48GB GPU for local serving
- Text-only (multimodal requires Qwen2-VL)
Capabilities & Modalities
Best Production Use Cases
- Multilingual enterprise customer service and translation
- Mathematical problem solving and scientific data modeling
- Self-hosted general LLM platform
Verified Benchmark Results
Standardized evaluations with methodology notes and authoritative citation links.
Pricing & Inference Cost Calculator
Token Cost Estimator – Qwen 2.5 72B Instruct
Calculate projected inference spend with prompt caching
Quick Workload Presets
$0.0000
2.00M input tokens
$0.0000
0.50M output tokens
$0.00
Avg: $0.00000 / req
Compare with Similar Models
DeepSeek-R1
Open-weights frontier reasoning model trained via large-scale reinforcement learning without supervised cold start.
Codestral 2501
Mistral AI's dedicated code generation model with 256k context window and fill-in-the-middle (FIM) capabilities.
DeepSeek-V3
Frontier open-weights 671B MoE base and chat model with multi-head latent attention (MLA) and dual-pipe training.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.