Reasoning & Math Evaluation Benchmark

AIME (2024/2025)

American Invitational Mathematics Examination

Prestigious high school mathematics competition problems requiring complex multi-step creative deduction and integer answers from 000 to 999.

Verified Model Leaderboard

Ranked by verified % Accuracy (Pass@1) score from official technical reports.

Primary Metric: % Accuracy (Pass@1)
RankAI Model & ProviderScore (% Accuracy (Pass@1))Evaluation MethodologyVerification DateSource
#1OpenAI o3-mini

OpenAI Reinforcement Learning Reasoning Model

87.3%Pass@1 with high reasoning effort (13.1/15 problems solved)Jan 2025OpenAI o3-mini Announcement
#2OpenAI o1

OpenAI Reinforcement Learning Reasoning Model

83.3%Pass@1 with high reasoning effort (12.5/15 problems solved)Dec 2024OpenAI o1 System Card
#3Claude 3.7 Sonnet

Anthropic Hybrid Reasoning Transformer

80.0%Pass@1 with 64k thinking budgetFeb 2025Anthropic Technical Report
#4DeepSeek-R1

DeepSeek MoE (Mixture of Experts) Reasoning

79.8%Pass@1 with native <think> chain-of-thought tokensJan 2025DeepSeek-R1 Technical Report
#5Gemini 2.0 Pro (Experimental)

Google DeepMind Multimodal Transformer

76.5%Pass@1 mathematical reasoningFeb 2025Google DeepMind Technical Report

Evaluation Methodology & Harness

15 questions per exam evaluated strictly without partial credit; evaluated using single-pass reasoning (Pass@1) or consensus sampling (Consensus@64).

Dataset & Test Set Construction

Official AIME 2024 and 2025 examination problems with exact integer answer verification.

Why This Benchmark Matters

The premier test for mathematical reasoning in frontier models, heavily used to measure the breakthrough of reasoning models like o1, o3-mini, and DeepSeek-R1.

Known Limitations & Contamination Risks

Only 15 questions per year, so standard practice aggregates across recent competition years or reports exact pass rates.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.