AIME (2024/2025)
American Invitational Mathematics Examination
Prestigious high school mathematics competition problems requiring complex multi-step creative deduction and integer answers from 000 to 999.
Verified Model Leaderboard
Ranked by verified % Accuracy (Pass@1) score from official technical reports.
| Rank | AI Model & Provider | Score (% Accuracy (Pass@1)) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | OpenAI o3-mini OpenAI • Reinforcement Learning Reasoning Model | 87.3% | Pass@1 with high reasoning effort (13.1/15 problems solved) | Jan 2025 | OpenAI o3-mini Announcement |
| #2 | OpenAI o1 OpenAI • Reinforcement Learning Reasoning Model | 83.3% | Pass@1 with high reasoning effort (12.5/15 problems solved) | Dec 2024 | OpenAI o1 System Card |
| #3 | Claude 3.7 Sonnet Anthropic • Hybrid Reasoning Transformer | 80.0% | Pass@1 with 64k thinking budget | Feb 2025 | Anthropic Technical Report |
| #4 | DeepSeek-R1 DeepSeek • MoE (Mixture of Experts) Reasoning | 79.8% | Pass@1 with native <think> chain-of-thought tokens | Jan 2025 | DeepSeek-R1 Technical Report |
| #5 | Gemini 2.0 Pro (Experimental) Google DeepMind • Multimodal Transformer | 76.5% | Pass@1 mathematical reasoning | Feb 2025 | Google DeepMind Technical Report |
Evaluation Methodology & Harness
15 questions per exam evaluated strictly without partial credit; evaluated using single-pass reasoning (Pass@1) or consensus sampling (Consensus@64).
Official AIME 2024 and 2025 examination problems with exact integer answer verification.
Why This Benchmark Matters
The premier test for mathematical reasoning in frontier models, heavily used to measure the breakthrough of reasoning models like o1, o3-mini, and DeepSeek-R1.
Known Limitations & Contamination Risks
Only 15 questions per year, so standard practice aggregates across recent competition years or reports exact pass rates.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.