MMMU
Massive Multi-discipline Multimodal Understanding
Multimodal benchmark assessing visual reasoning across 30 college-level subjects with diagrams, charts, medical scans, maps, and technical illustrations.
Verified Model Leaderboard
Ranked by verified % Accuracy score from official technical reports.
| Rank | AI Model & Provider | Score (% Accuracy) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | Gemini 2.0 Pro (Experimental) Google DeepMind • Multimodal Transformer | 76.2% | Multimodal visual benchmark | Feb 2025 | Google DeepMind Technical Report |
| #2 | Gemini 2.0 Flash Google DeepMind • Multimodal MoE | 73.1% | Multimodal visual benchmark | Dec 2024 | Google DeepMind Report |
| #3 | Claude 3.7 Sonnet Anthropic • Hybrid Reasoning Transformer | 72.5% | Pass@1 multimodal visual reasoning | Feb 2025 | Anthropic Model Card |
| #4 | Claude 3.5 Sonnet (v2) Anthropic • Dense Transformer | 70.4% | Multimodal college level visual evaluation | Oct 2024 | Anthropic Model Card |
| #5 | GPT-4o OpenAI • Multimodal Transformer | 69.1% | Multimodal visual benchmark | May 2024 | OpenAI GPT-4o Card |
| #6 | GPT-4o mini OpenAI • Multimodal Small Transformer | 59.4% | Multimodal visual reasoning benchmark | Jul 2024 | OpenAI Model Card |
Evaluation Methodology & Harness
Questions require joint comprehension of complex domain text and technical imagery, evaluated with standard multiple choice and open-ended answer verification.
11,500 questions across 30 subjects requiring college-level expert multimodal reasoning.
Why This Benchmark Matters
The standard benchmark for evaluating multimodal frontier models (GPT-4o, Claude 3.7 Sonnet, Gemini 2.0 Flash) on professional visual tasks.
Known Limitations & Contamination Risks
Images require high-resolution encoding; OCR capabilities can introduce confounding variables.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.