Reasoning & Math Evaluation Benchmark

ARC-AGI

Abstraction and Reasoning Corpus for AGI

Benchmark created by François Chollet measuring an AI system's ability to acquire new skills and solve novel visual grid puzzles with minimal priors.

Verified Model Leaderboard

Ranked by verified % Pass@2 score from official technical reports.

Primary Metric: % Pass@2
RankAI Model & ProviderScore (% Pass@2)Evaluation MethodologyVerification DateSource
#1OpenAI o1

OpenAI Reinforcement Learning Reasoning Model

75.7%Pass@2 with test-time search scaffoldingDec 2024ARC Prize Leaderboard / OpenAI

Evaluation Methodology & Harness

Each task features 3-5 demonstration input/output grid pairs and requires predicting the test output grid without task-specific pretraining.

Dataset & Test Set Construction

400 public training tasks and 400 private evaluation tasks designed to measure efficient broad skill acquisition.

Why This Benchmark Matters

Tests general fluid intelligence and novel reasoning rather than knowledge memorization.

Known Limitations & Contamination Risks

Extremely difficult for standard LLMs without test-time compute scaling or programmatic search scaffolding.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.