General Knowledge & Frontier Evaluation Benchmark

Humanity's Last Exam (HLE)

Humanity's Last Exam: A Benchmark for Frontier AI

A rigorous multimodal benchmark developed by CAIS and Scale AI designed to be resistant to saturation, consisting of 3,000 expert-level questions spanning dozens of academic and professional fields.

Verified Model Leaderboard

Ranked by verified % Accuracy score from official technical reports.

Primary Metric: % Accuracy
RankAI Model & ProviderScore (% Accuracy)Evaluation MethodologyVerification DateSource
#1Claude 3.7 Sonnet

Anthropic Hybrid Reasoning Transformer

22.8%Zero-shot with thinking budget on CAIS Humanity's Last ExamFeb 2025CAIS / Anthropic Technical Evaluation
#2OpenAI o1

OpenAI Reinforcement Learning Reasoning Model

19.5%Zero-shot with high reasoning tokens on CAIS benchmarkDec 2024CAIS Humanity's Last Exam Report

Evaluation Methodology & Harness

Questions are closed-ended, multiple choice or exact match, reviewed by subject-matter PhDs. Evaluated in zero-shot or few-shot settings with strict contamination filters.

Dataset & Test Set Construction

3,000 multi-discipline questions written by global subject-matter experts across humanities, STEM, medicine, and law.

Why This Benchmark Matters

Acts as the ultimate frontier benchmark to separate current reasoning models from true domain mastery where standard benchmarks like MMLU are saturated (>90%).

Known Limitations & Contamination Risks

High difficulty means current frontier models score below 20-30%, resulting in high variance across small question subsets.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.