LiveCodeBench
LiveCodeBench: Holistic and Contamination-Free Code Evaluation
Continuously updated programming benchmark collecting problems published after model training cutoffs from LeetCode, AtCoder, and Codeforces to prevent data contamination.
Verified Model Leaderboard
Ranked by verified Pass@1 (%) score from official technical reports.
| Rank | AI Model & Provider | Score (Pass@1 (%)) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | Claude 3.7 Sonnet Anthropic • Hybrid Reasoning Transformer | 70.3% | Pass@1, 0-shot with extended thinking budget | Feb 2025 | LiveCodeBench / Anthropic Evaluation |
| #2 | OpenAI o3-mini OpenAI • Reinforcement Learning Reasoning Model | 68.2% | Pass@1 with high reasoning effort | Jan 2025 | LiveCodeBench Evaluation |
| #3 | DeepSeek-R1 DeepSeek • MoE (Mixture of Experts) Reasoning | 65.9% | Pass@1 0-shot with reasoning tokens | Jan 2025 | LiveCodeBench Official Leaderboard |
| #4 | Qwen 2.5 Coder 32B Instruct Qwen (Alibaba Cloud) • Dense Transformer | 60.1% | Pass@1, 0-shot code generation | Nov 2024 | LiveCodeBench Official Leaderboard |
| #5 | Gemini 2.0 Pro (Experimental) Google DeepMind • Multimodal Transformer | 58.4% | Pass@1, 0-shot code generation | Feb 2025 | LiveCodeBench Leaderboard |
| #6 | Claude 3.5 Sonnet (v2) Anthropic • Dense Transformer | 52.4% | Pass@1, 0-shot code generation | Oct 2024 | LiveCodeBench Official Leaderboard |
| #7 | DeepSeek-V3 DeepSeek • MoE (Multi-Head Latent Attention) | 49.2% | Pass@1, 0-shot code generation | Dec 2024 | LiveCodeBench Leaderboard |
| #8 | Gemini 2.0 Flash Google DeepMind • Multimodal MoE | 48.0% | Pass@1, 0-shot code generation | Dec 2024 | LiveCodeBench Leaderboard |
| #9 | Llama 3.3 70B Instruct Meta AI • Dense Transformer | 47.9% | Pass@1, 0-shot code generation | Dec 2024 | LiveCodeBench Leaderboard |
| #10 | GPT-4o OpenAI • Multimodal Transformer | 45.3% | Pass@1, 0-shot code generation | Aug 2024 | LiveCodeBench Official Leaderboard |
| #11 | Claude 3.5 Haiku Anthropic • Dense Transformer | 43.1% | Pass@1, 0-shot code generation | Nov 2024 | LiveCodeBench Leaderboard |
| #12 | GPT-4o mini OpenAI • Multimodal Small Transformer | 32.4% | Pass@1, 0-shot | Jul 2024 | LiveCodeBench Leaderboard |
Evaluation Methodology & Harness
Problems are categorized into easy, medium, and hard. Evaluated with Pass@1 across standard input/output test suites without prior exposure in training corpora.
Over 500 competitive programming challenges continuously synced from live competitions (2024-2025).
Why This Benchmark Matters
Eliminates memorization biases present in older benchmarks like HumanEval; tests algorithmic problem-solving and exact code generation.
Known Limitations & Contamination Risks
Primarily tests competitive programming algorithms rather than multi-file system architecture.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.