Coding Evaluation Benchmark

LiveCodeBench Pro

LiveCodeBench Pro: Frontier Hard Algorithmic Benchmark

Hard subset of LiveCodeBench focusing on advanced competitive programming problems (Div 1 / Div 2 Codeforces and Hard LeetCode).

Verified Model Leaderboard

Ranked by verified Pass@1 (%) score from official technical reports.

Primary Metric: Pass@1 (%)
RankAI Model & ProviderScore (Pass@1 (%))Evaluation MethodologyVerification DateSource
#1OpenAI o1

OpenAI Reinforcement Learning Reasoning Model

63.8%Pass@1 on hard competitive problemsDec 2024LiveCodeBench Evaluation

Evaluation Methodology & Harness

Pass@1 evaluation with strict execution time limits and memory boundaries to measure frontier algorithmic reasoning.

Dataset & Test Set Construction

Curated set of 200+ Hard-tier competitive problems with comprehensive stress test suites.

Why This Benchmark Matters

Differentiates reasoning models (o1, o3-mini, DeepSeek-R1, Claude 3.7 Sonnet) from standard instruction-tuned models.

Known Limitations & Contamination Risks

Extremely challenging; standard non-reasoning models often score below 25%.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.