AI Benchmarks
In-depth methodology documentation, test harness analysis, and verified performance leaderboards for modern frontier AI evaluation.
CursorBench
CursorBench: Real-World IDE Code Editing and Navigation
Practical benchmark measuring an AI model's ability to perform code edits, fast diff generation, symbol resolution, and context retrieval in realistic editor workflows.
SWE-bench Verified
SWE-bench: Resolving Real-World GitHub Issues (Verified Subset)
Standard benchmark evaluating an AI model's ability to solve end-to-end software engineering problems from real GitHub issues in popular Python repositories.
LiveCodeBench
LiveCodeBench: Holistic and Contamination-Free Code Evaluation
Continuously updated programming benchmark collecting problems published after model training cutoffs from LeetCode, AtCoder, and Codeforces to prevent data contamination.
LiveCodeBench Pro
LiveCodeBench Pro: Frontier Hard Algorithmic Benchmark
Hard subset of LiveCodeBench focusing on advanced competitive programming problems (Div 1 / Div 2 Codeforces and Hard LeetCode).
Humanity's Last Exam (HLE)
Humanity's Last Exam: A Benchmark for Frontier AI
A rigorous multimodal benchmark developed by CAIS and Scale AI designed to be resistant to saturation, consisting of 3,000 expert-level questions spanning dozens of academic and professional fields.
MMLU-Pro
MMLU-Pro: A More Robust and Challenging Multi-Task Benchmark
An enhanced version of Massive Multitask Language Understanding (MMLU) featuring 10-choice questions and a higher proportion of reasoning-focused problems.
MMMU
Massive Multi-discipline Multimodal Understanding
Multimodal benchmark assessing visual reasoning across 30 college-level subjects with diagrams, charts, medical scans, maps, and technical illustrations.
GPQA Diamond
Google-Proof Q&A (Diamond Subset)
Graduate-level physics, chemistry, and biology questions written by domain experts. Non-expert humans with unrestricted web access score only 34%.
AIME (2024/2025)
American Invitational Mathematics Examination
Prestigious high school mathematics competition problems requiring complex multi-step creative deduction and integer answers from 000 to 999.
ARC-AGI
Abstraction and Reasoning Corpus for AGI
Benchmark created by François Chollet measuring an AI system's ability to acquire new skills and solve novel visual grid puzzles with minimal priors.
BFCL (Berkeley Function Calling)
Berkeley Function Calling Leaderboard
Comprehensive evaluation of LLM tool-calling and function-calling capabilities across simple, multi-turn, parallel, and executable API scenarios.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.