Coding Evaluation Benchmark

LiveCodeBench

LiveCodeBench: Holistic and Contamination-Free Code Evaluation

Continuously updated programming benchmark collecting problems published after model training cutoffs from LeetCode, AtCoder, and Codeforces to prevent data contamination.

Verified Model Leaderboard

Ranked by verified Pass@1 (%) score from official technical reports.

Primary Metric: Pass@1 (%)
RankAI Model & ProviderScore (Pass@1 (%))Evaluation MethodologyVerification DateSource
#1Claude 3.7 Sonnet

Anthropic Hybrid Reasoning Transformer

70.3%Pass@1, 0-shot with extended thinking budgetFeb 2025LiveCodeBench / Anthropic Evaluation
#2OpenAI o3-mini

OpenAI Reinforcement Learning Reasoning Model

68.2%Pass@1 with high reasoning effortJan 2025LiveCodeBench Evaluation
#3DeepSeek-R1

DeepSeek MoE (Mixture of Experts) Reasoning

65.9%Pass@1 0-shot with reasoning tokensJan 2025LiveCodeBench Official Leaderboard
#4Qwen 2.5 Coder 32B Instruct

Qwen (Alibaba Cloud) Dense Transformer

60.1%Pass@1, 0-shot code generationNov 2024LiveCodeBench Official Leaderboard
#5Gemini 2.0 Pro (Experimental)

Google DeepMind Multimodal Transformer

58.4%Pass@1, 0-shot code generationFeb 2025LiveCodeBench Leaderboard
#6Claude 3.5 Sonnet (v2)

Anthropic Dense Transformer

52.4%Pass@1, 0-shot code generationOct 2024LiveCodeBench Official Leaderboard
#7DeepSeek-V3

DeepSeek MoE (Multi-Head Latent Attention)

49.2%Pass@1, 0-shot code generationDec 2024LiveCodeBench Leaderboard
#8Gemini 2.0 Flash

Google DeepMind Multimodal MoE

48.0%Pass@1, 0-shot code generationDec 2024LiveCodeBench Leaderboard
#9Llama 3.3 70B Instruct

Meta AI Dense Transformer

47.9%Pass@1, 0-shot code generationDec 2024LiveCodeBench Leaderboard
#10GPT-4o

OpenAI Multimodal Transformer

45.3%Pass@1, 0-shot code generationAug 2024LiveCodeBench Official Leaderboard
#11Claude 3.5 Haiku

Anthropic Dense Transformer

43.1%Pass@1, 0-shot code generationNov 2024LiveCodeBench Leaderboard
#12GPT-4o mini

OpenAI Multimodal Small Transformer

32.4%Pass@1, 0-shotJul 2024LiveCodeBench Leaderboard

Evaluation Methodology & Harness

Problems are categorized into easy, medium, and hard. Evaluated with Pass@1 across standard input/output test suites without prior exposure in training corpora.

Dataset & Test Set Construction

Over 500 competitive programming challenges continuously synced from live competitions (2024-2025).

Why This Benchmark Matters

Eliminates memorization biases present in older benchmarks like HumanEval; tests algorithmic problem-solving and exact code generation.

Known Limitations & Contamination Risks

Primarily tests competitive programming algorithms rather than multi-file system architecture.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.