MMLU-Pro
MMLU-Pro: A More Robust and Challenging Multi-Task Benchmark
An enhanced version of Massive Multitask Language Understanding (MMLU) featuring 10-choice questions and a higher proportion of reasoning-focused problems.
Verified Model Leaderboard
Ranked by verified % Accuracy score from official technical reports.
| Rank | AI Model & Provider | Score (% Accuracy) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | DeepSeek-R1 DeepSeek • MoE (Mixture of Experts) Reasoning | 84.0% | 5-shot with full reasoning chain | Jan 2025 | DeepSeek-R1 Technical Report |
| #2 | Claude 3.5 Sonnet (v2) Anthropic • Dense Transformer | 78.0% | 5-shot standard Chain-of-Thought prompt | Oct 2024 | TIGER-AI-Lab MMLU-Pro Leaderboard |
| #3 | DeepSeek-V3 DeepSeek • MoE (Multi-Head Latent Attention) | 75.9% | 5-shot standard evaluation | Dec 2024 | DeepSeek-V3 Technical Report |
| #4 | Gemini 2.0 Flash Google DeepMind • Multimodal MoE | 74.8% | 5-shot Chain-of-Thought standard | Dec 2024 | Google DeepMind Report |
| #5 | GPT-4o OpenAI • Multimodal Transformer | 72.5% | 5-shot standard benchmark run | May 2024 | MMLU-Pro Leaderboard |
| #6 | Llama 3.3 70B Instruct Meta AI • Dense Transformer | 71.0% | 5-shot CoT evaluation | Dec 2024 | Meta AI Model Card |
| #7 | Qwen 2.5 Coder 32B Instruct Qwen (Alibaba Cloud) • Dense Transformer | 70.4% | 5-shot standard evaluation | Nov 2024 | Qwen 2.5 Coder Technical Report |
| #8 | Claude 3.5 Haiku Anthropic • Dense Transformer | 65.1% | 5-shot Chain-of-Thought standard | Nov 2024 | Anthropic Model Card |
| #9 | GPT-4o mini OpenAI • Multimodal Small Transformer | 63.0% | 5-shot CoT evaluation | Jul 2024 | MMLU-Pro Leaderboard |
Evaluation Methodology & Harness
12,000+ questions across 14 subjects with 10 answer options instead of 4, dramatically reducing random guessing and measuring true reasoning depth.
12,000+ college and graduate-level problems in biology, physics, chemistry, computer science, law, and economics.
Why This Benchmark Matters
Provides a much more discriminative evaluation than classic MMLU, which has reached saturation for top frontier models.
Known Limitations & Contamination Risks
Requires strong zero-shot / 5-shot CoT prompting to assess reasoning properly.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.