Coding Evaluation Benchmark

CursorBench

CursorBench: Real-World IDE Code Editing and Navigation

Practical benchmark measuring an AI model's ability to perform code edits, fast diff generation, symbol resolution, and context retrieval in realistic editor workflows.

Verified Model Leaderboard

Ranked by verified Score (%) score from official technical reports.

Primary Metric: Score (%)
Leaderboard scores for this benchmark are currently being compiled.

Evaluation Methodology & Harness

Evaluates structured edit diff accuracy, adherence to repository style, and multi-file contextual updates across production codebases.

Dataset & Test Set Construction

Hundreds of realistic code modification scenarios and refactoring tasks in active open-source projects.

Why This Benchmark Matters

Directly reflects day-to-day developer productivity in AI code editors like Cursor, VS Code, and Windsurf.

Known Limitations & Contamination Risks

Fast-moving benchmark dependent on editor client orchestration harness.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.