Multimodal Evaluation Benchmark

MMMU

Massive Multi-discipline Multimodal Understanding

Multimodal benchmark assessing visual reasoning across 30 college-level subjects with diagrams, charts, medical scans, maps, and technical illustrations.

Verified Model Leaderboard

Ranked by verified % Accuracy score from official technical reports.

Primary Metric: % Accuracy
RankAI Model & ProviderScore (% Accuracy)Evaluation MethodologyVerification DateSource
#1Gemini 2.0 Pro (Experimental)

Google DeepMind Multimodal Transformer

76.2%Multimodal visual benchmarkFeb 2025Google DeepMind Technical Report
#2Gemini 2.0 Flash

Google DeepMind Multimodal MoE

73.1%Multimodal visual benchmarkDec 2024Google DeepMind Report
#3Claude 3.7 Sonnet

Anthropic Hybrid Reasoning Transformer

72.5%Pass@1 multimodal visual reasoningFeb 2025Anthropic Model Card
#4Claude 3.5 Sonnet (v2)

Anthropic Dense Transformer

70.4%Multimodal college level visual evaluationOct 2024Anthropic Model Card
#5GPT-4o

OpenAI Multimodal Transformer

69.1%Multimodal visual benchmarkMay 2024OpenAI GPT-4o Card
#6GPT-4o mini

OpenAI Multimodal Small Transformer

59.4%Multimodal visual reasoning benchmarkJul 2024OpenAI Model Card

Evaluation Methodology & Harness

Questions require joint comprehension of complex domain text and technical imagery, evaluated with standard multiple choice and open-ended answer verification.

Dataset & Test Set Construction

11,500 questions across 30 subjects requiring college-level expert multimodal reasoning.

Why This Benchmark Matters

The standard benchmark for evaluating multimodal frontier models (GPT-4o, Claude 3.7 Sonnet, Gemini 2.0 Flash) on professional visual tasks.

Known Limitations & Contamination Risks

Images require high-resolution encoding; OCR capabilities can introduce confounding variables.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.