ARC-AGI
Abstraction and Reasoning Corpus for AGI
Benchmark created by François Chollet measuring an AI system's ability to acquire new skills and solve novel visual grid puzzles with minimal priors.
Verified Model Leaderboard
Ranked by verified % Pass@2 score from official technical reports.
| Rank | AI Model & Provider | Score (% Pass@2) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | OpenAI o1 OpenAI • Reinforcement Learning Reasoning Model | 75.7% | Pass@2 with test-time search scaffolding | Dec 2024 | ARC Prize Leaderboard / OpenAI |
Evaluation Methodology & Harness
Each task features 3-5 demonstration input/output grid pairs and requires predicting the test output grid without task-specific pretraining.
400 public training tasks and 400 private evaluation tasks designed to measure efficient broad skill acquisition.
Why This Benchmark Matters
Tests general fluid intelligence and novel reasoning rather than knowledge memorization.
Known Limitations & Contamination Risks
Extremely difficult for standard LLMs without test-time compute scaling or programmatic search scaffolding.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.