Standards & Evaluation
Benchmark Methodology
Our framework for rigorous, contamination-aware, and reproducible AI evaluation
Published: July 2026
1. Core Evaluation Philosophy
Our benchmark intelligence system is designed to provide transparent, scientifically grounded comparisons of foundation models. We reject saturated, legacy benchmarks in favor of dynamic, contamination-free, and high-difficulty evaluations such as Humanity's Last Exam, SWE-bench Verified, LiveCodeBench, and AIME.
We strictly uphold the principle that benchmark numbers are only meaningful when paired with their exact prompting methodology, few-shot samples, and pass criteria.
2. Standards for Comparability
We never present scores as directly comparable when the evaluation harness or test conditions differ.
Key methodological factors tracked in our database include:
• Zero-Shot vs. Few-Shot Prompting (e.g. 0-shot CoT vs. 5-shot standard).
• Pass Criteria (e.g. Pass@1 single generation vs. Consensus@64 majority voting).
• Test-Time Compute & Thinking Budgets (e.g. reasoning token allocations on o1/o3-mini/Claude 3.7 Sonnet).
• Agentic Scaffolding vs. Raw Model Outputs.
3. Data Contamination Safeguards
Data contamination—where benchmark problems inadvertently leak into a model's pretraining corpora—is a primary concern in AI evaluation.
We actively highlight benchmarks with active contamination protections, such as LiveCodeBench (which continuously sources problems published after model cutoff dates) and private evaluation test sets.
4. Authoritative Sourcing Hierarchy
Every benchmark score displayed in our database is attributed according to a strict priority hierarchy:
1. Official peer-reviewed technical reports and system cards from primary research institutions.
2. Official benchmark maintainer leaderboards (e.g. SWE-bench Official, LiveCodeBench, CAIS, Berkeley Gorilla BFCL).
3. Independent academic evaluations with publicly auditable execution logs.