Humanity's Last Exam (HLE)
Humanity's Last Exam: A Benchmark for Frontier AI
A rigorous multimodal benchmark developed by CAIS and Scale AI designed to be resistant to saturation, consisting of 3,000 expert-level questions spanning dozens of academic and professional fields.
Verified Model Leaderboard
Ranked by verified % Accuracy score from official technical reports.
| Rank | AI Model & Provider | Score (% Accuracy) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | Claude 3.7 Sonnet Anthropic • Hybrid Reasoning Transformer | 22.8% | Zero-shot with thinking budget on CAIS Humanity's Last Exam | Feb 2025 | CAIS / Anthropic Technical Evaluation |
| #2 | OpenAI o1 OpenAI • Reinforcement Learning Reasoning Model | 19.5% | Zero-shot with high reasoning tokens on CAIS benchmark | Dec 2024 | CAIS Humanity's Last Exam Report |
Evaluation Methodology & Harness
Questions are closed-ended, multiple choice or exact match, reviewed by subject-matter PhDs. Evaluated in zero-shot or few-shot settings with strict contamination filters.
3,000 multi-discipline questions written by global subject-matter experts across humanities, STEM, medicine, and law.
Why This Benchmark Matters
Acts as the ultimate frontier benchmark to separate current reasoning models from true domain mastery where standard benchmarks like MMLU are saturated (>90%).
Known Limitations & Contamination Risks
High difficulty means current frontier models score below 20-30%, resulting in high variance across small question subsets.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.