General Knowledge & Frontier Evaluation Benchmark

MMLU-Pro

MMLU-Pro: A More Robust and Challenging Multi-Task Benchmark

An enhanced version of Massive Multitask Language Understanding (MMLU) featuring 10-choice questions and a higher proportion of reasoning-focused problems.

Verified Model Leaderboard

Ranked by verified % Accuracy score from official technical reports.

Primary Metric: % Accuracy
RankAI Model & ProviderScore (% Accuracy)Evaluation MethodologyVerification DateSource
#1DeepSeek-R1

DeepSeek MoE (Mixture of Experts) Reasoning

84.0%5-shot with full reasoning chainJan 2025DeepSeek-R1 Technical Report
#2Claude 3.5 Sonnet (v2)

Anthropic Dense Transformer

78.0%5-shot standard Chain-of-Thought promptOct 2024TIGER-AI-Lab MMLU-Pro Leaderboard
#3DeepSeek-V3

DeepSeek MoE (Multi-Head Latent Attention)

75.9%5-shot standard evaluationDec 2024DeepSeek-V3 Technical Report
#4Gemini 2.0 Flash

Google DeepMind Multimodal MoE

74.8%5-shot Chain-of-Thought standardDec 2024Google DeepMind Report
#5GPT-4o

OpenAI Multimodal Transformer

72.5%5-shot standard benchmark runMay 2024MMLU-Pro Leaderboard
#6Llama 3.3 70B Instruct

Meta AI Dense Transformer

71.0%5-shot CoT evaluationDec 2024Meta AI Model Card
#7Qwen 2.5 Coder 32B Instruct

Qwen (Alibaba Cloud) Dense Transformer

70.4%5-shot standard evaluationNov 2024Qwen 2.5 Coder Technical Report
#8Claude 3.5 Haiku

Anthropic Dense Transformer

65.1%5-shot Chain-of-Thought standardNov 2024Anthropic Model Card
#9GPT-4o mini

OpenAI Multimodal Small Transformer

63.0%5-shot CoT evaluationJul 2024MMLU-Pro Leaderboard

Evaluation Methodology & Harness

12,000+ questions across 14 subjects with 10 answer options instead of 4, dramatically reducing random guessing and measuring true reasoning depth.

Dataset & Test Set Construction

12,000+ college and graduate-level problems in biology, physics, chemistry, computer science, law, and economics.

Why This Benchmark Matters

Provides a much more discriminative evaluation than classic MMLU, which has reached saturation for top frontier models.

Known Limitations & Contamination Risks

Requires strong zero-shot / 5-shot CoT prompting to assess reasoning properly.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.