Standardized Evaluation Frameworks & Leaderboards

AI Benchmarks

In-depth methodology documentation, test harness analysis, and verified performance leaderboards for modern frontier AI evaluation.

CodingMetric: Score (%)

CursorBench

CursorBench: Real-World IDE Code Editing and Navigation

Practical benchmark measuring an AI model's ability to perform code edits, fast diff generation, symbol resolution, and context retrieval in realistic editor workflows.

0 verified model evaluationsLeaderboard & Specs
CodingMetric: Pass@1 (%)

SWE-bench Verified

SWE-bench: Resolving Real-World GitHub Issues (Verified Subset)

Standard benchmark evaluating an AI model's ability to solve end-to-end software engineering problems from real GitHub issues in popular Python repositories.

9 verified model evaluationsLeaderboard & Specs
CodingMetric: Pass@1 (%)

LiveCodeBench

LiveCodeBench: Holistic and Contamination-Free Code Evaluation

Continuously updated programming benchmark collecting problems published after model training cutoffs from LeetCode, AtCoder, and Codeforces to prevent data contamination.

12 verified model evaluationsLeaderboard & Specs
CodingMetric: Pass@1 (%)

LiveCodeBench Pro

LiveCodeBench Pro: Frontier Hard Algorithmic Benchmark

Hard subset of LiveCodeBench focusing on advanced competitive programming problems (Div 1 / Div 2 Codeforces and Hard LeetCode).

1 verified model evaluationsLeaderboard & Specs
General Knowledge & FrontierMetric: % Accuracy

Humanity's Last Exam (HLE)

Humanity's Last Exam: A Benchmark for Frontier AI

A rigorous multimodal benchmark developed by CAIS and Scale AI designed to be resistant to saturation, consisting of 3,000 expert-level questions spanning dozens of academic and professional fields.

2 verified model evaluationsLeaderboard & Specs
General Knowledge & FrontierMetric: % Accuracy

MMLU-Pro

MMLU-Pro: A More Robust and Challenging Multi-Task Benchmark

An enhanced version of Massive Multitask Language Understanding (MMLU) featuring 10-choice questions and a higher proportion of reasoning-focused problems.

9 verified model evaluationsLeaderboard & Specs
MultimodalMetric: % Accuracy

MMMU

Massive Multi-discipline Multimodal Understanding

Multimodal benchmark assessing visual reasoning across 30 college-level subjects with diagrams, charts, medical scans, maps, and technical illustrations.

6 verified model evaluationsLeaderboard & Specs
Reasoning & MathMetric: % Accuracy

GPQA Diamond

Google-Proof Q&A (Diamond Subset)

Graduate-level physics, chemistry, and biology questions written by domain experts. Non-expert humans with unrestricted web access score only 34%.

12 verified model evaluationsLeaderboard & Specs
Reasoning & MathMetric: % Accuracy (Pass@1)

AIME (2024/2025)

American Invitational Mathematics Examination

Prestigious high school mathematics competition problems requiring complex multi-step creative deduction and integer answers from 000 to 999.

5 verified model evaluationsLeaderboard & Specs
Reasoning & MathMetric: % Pass@2

ARC-AGI

Abstraction and Reasoning Corpus for AGI

Benchmark created by François Chollet measuring an AI system's ability to acquire new skills and solve novel visual grid puzzles with minimal priors.

1 verified model evaluationsLeaderboard & Specs
Tool & AgenticMetric: % AST Accuracy

BFCL (Berkeley Function Calling)

Berkeley Function Calling Leaderboard

Comprehensive evaluation of LLM tool-calling and function-calling capabilities across simple, multi-turn, parallel, and executable API scenarios.

6 verified model evaluationsLeaderboard & Specs

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.