Qwen (Alibaba Cloud) Open Weights (Apache 2.0)Status: active

Qwen 2.5 72B Instruct

Alibaba's flagship open foundation model with world-class multilingual, mathematical, and coding capabilities.

Released: Sep 19, 2024
Last Verified: Aug 10, 2026
Context Window

128k

Tokens

Max Output

8.192k

Output limit

Architecture

Dense Transformer

Model family

Parameters

72.7B

Total / Active

Input Token Cost

$0.00

Per 1M tokens

Output Token Cost

$0.00

Per 1M tokens

Model Overview

Qwen 2.5 72B Instruct is an open-weights dense transformer trained on 18 trillion tokens with superior mathematical reasoning, 128k context length, and deep multilingual competence across 29+ languages.

Developer Implementation Notes

Apache 2.0 license. Scores 83.1% on MMLU-Pro and 85.5% on MATH-500.

Key Strengths

  • Apache 2.0 permissive license
  • Exceptional mathematical and reasoning benchmark scores
  • Support for 29+ natural languages
  • 128k context window

Limitations & Boundaries

  • Requires dual 24GB GPUs or single 48GB GPU for local serving
  • Text-only (multimodal requires Qwen2-VL)

Capabilities & Modalities

Modalities:text
Reasoning Chains:No
Tool / Function Calling:Supported
Structured Outputs:Supported
Fine-Tuning:Supported
Local Deployment:Yes
Min VRAM:38 GB

Best Production Use Cases

  • Multilingual enterprise customer service and translation
  • Mathematical problem solving and scientific data modeling
  • Self-hosted general LLM platform

Verified Benchmark Results

Standardized evaluations with methodology notes and authoritative citation links.

View all industry benchmarks →
Official benchmark results are currently being verified for this model release.

Pricing & Inference Cost Calculator

Token Cost Estimator – Qwen 2.5 72B Instruct

Calculate projected inference spend with prompt caching

Quick Workload Presets

Input Tokens / Req2,000
100100k200k
Output Tokens / Req500
5016k32k
Number of Requests1,000
125k50k
Cache Hit Rate0%
0% (No cache)50%90% (Max)
Total Input Spend

$0.0000

2.00M input tokens

Total Output Spend

$0.0000

0.50M output tokens

Estimated Total Cost

$0.00

Avg: $0.00000 / req

Rates: $0.00 in / $0.00 out per million tokens (Self-Hosted (Local))Last verified: Aug 10, 2026

Compare with Similar Models

DeepSeek

DeepSeek-R1

Open-weights frontier reasoning model trained via large-scale reinforcement learning without supervised cold start.

Compare Qwen 2.5 72B Instruct vs DeepSeek-R1
Mistral AI

Codestral 2501

Mistral AI's dedicated code generation model with 256k context window and fill-in-the-middle (FIM) capabilities.

Compare Qwen 2.5 72B Instruct vs Codestral 2501
DeepSeek

DeepSeek-V3

Frontier open-weights 671B MoE base and chat model with multi-head latent attention (MLA) and dual-pipe training.

Compare Qwen 2.5 72B Instruct vs DeepSeek-V3

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.