Reasoning & Math Evaluation Benchmark

GPQA Diamond

Google-Proof Q&A (Diamond Subset)

Graduate-level physics, chemistry, and biology questions written by domain experts. Non-expert humans with unrestricted web access score only 34%.

Verified Model Leaderboard

Ranked by verified % Accuracy score from official technical reports.

Primary Metric: % Accuracy
RankAI Model & ProviderScore (% Accuracy)Evaluation MethodologyVerification DateSource
#1Claude 3.7 Sonnet

Anthropic Hybrid Reasoning Transformer

78.4%Zero-shot with Chain-of-Thought thinking budgetFeb 2025Anthropic Technical Report
#2OpenAI o3-mini

OpenAI Reinforcement Learning Reasoning Model

77.0%Zero-shot with high reasoning effortJan 2025OpenAI o3-mini Announcement
#3OpenAI o1

OpenAI Reinforcement Learning Reasoning Model

75.7%Zero-shot with reasoning tokensDec 2024OpenAI o1 System Card
#4Gemini 2.0 Pro (Experimental)

Google DeepMind Multimodal Transformer

74.2%Zero-shot Chain-of-ThoughtFeb 2025Google DeepMind Technical Report
#5DeepSeek-R1

DeepSeek MoE (Mixture of Experts) Reasoning

71.5%Zero-shot with reasoning tokensJan 2025DeepSeek-R1 Technical Report
#6Claude 3.5 Sonnet (v2)

Anthropic Dense Transformer

65.0%Zero-shot Chain-of-Thought promptingOct 2024Anthropic Technical Report
#7Gemini 2.0 Flash

Google DeepMind Multimodal MoE

62.1%Zero-shot Chain-of-Thought promptDec 2024Google DeepMind Report
#8DeepSeek-V3

DeepSeek MoE (Multi-Head Latent Attention)

59.1%Zero-shot Chain-of-ThoughtDec 2024DeepSeek-V3 Technical Report
#9GPT-4o

OpenAI Multimodal Transformer

53.6%Zero-shot Chain-of-ThoughtMay 2024OpenAI GPT-4o Card
#10Llama 3.3 70B Instruct

Meta AI Dense Transformer

52.8%Zero-shot Chain-of-ThoughtDec 2024Meta AI Model Card
#11Claude 3.5 Haiku

Anthropic Dense Transformer

41.6%Zero-shot Chain-of-ThoughtNov 2024Anthropic Model Card
#12GPT-4o mini

OpenAI Multimodal Small Transformer

40.2%Zero-shot Chain-of-ThoughtJul 2024OpenAI Model Card

Evaluation Methodology & Harness

Diamond subset comprises the highest-confidence, peer-validated 198 questions. Evaluated zero-shot with Chain-of-Thought (CoT).

Dataset & Test Set Construction

198 rigorously vetted graduate-level scientific problems requiring multi-step domain deduction.

Why This Benchmark Matters

Measures deep scientific reasoning and resilience against surface-level web search shortcuts.

Known Limitations & Contamination Risks

Small test set (198 questions) means statistical margins of error require careful reporting.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.