GPQA Diamond
Google-Proof Q&A (Diamond Subset)
Graduate-level physics, chemistry, and biology questions written by domain experts. Non-expert humans with unrestricted web access score only 34%.
Verified Model Leaderboard
Ranked by verified % Accuracy score from official technical reports.
| Rank | AI Model & Provider | Score (% Accuracy) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | Claude 3.7 Sonnet Anthropic • Hybrid Reasoning Transformer | 78.4% | Zero-shot with Chain-of-Thought thinking budget | Feb 2025 | Anthropic Technical Report |
| #2 | OpenAI o3-mini OpenAI • Reinforcement Learning Reasoning Model | 77.0% | Zero-shot with high reasoning effort | Jan 2025 | OpenAI o3-mini Announcement |
| #3 | OpenAI o1 OpenAI • Reinforcement Learning Reasoning Model | 75.7% | Zero-shot with reasoning tokens | Dec 2024 | OpenAI o1 System Card |
| #4 | Gemini 2.0 Pro (Experimental) Google DeepMind • Multimodal Transformer | 74.2% | Zero-shot Chain-of-Thought | Feb 2025 | Google DeepMind Technical Report |
| #5 | DeepSeek-R1 DeepSeek • MoE (Mixture of Experts) Reasoning | 71.5% | Zero-shot with reasoning tokens | Jan 2025 | DeepSeek-R1 Technical Report |
| #6 | Claude 3.5 Sonnet (v2) Anthropic • Dense Transformer | 65.0% | Zero-shot Chain-of-Thought prompting | Oct 2024 | Anthropic Technical Report |
| #7 | Gemini 2.0 Flash Google DeepMind • Multimodal MoE | 62.1% | Zero-shot Chain-of-Thought prompt | Dec 2024 | Google DeepMind Report |
| #8 | DeepSeek-V3 DeepSeek • MoE (Multi-Head Latent Attention) | 59.1% | Zero-shot Chain-of-Thought | Dec 2024 | DeepSeek-V3 Technical Report |
| #9 | GPT-4o OpenAI • Multimodal Transformer | 53.6% | Zero-shot Chain-of-Thought | May 2024 | OpenAI GPT-4o Card |
| #10 | Llama 3.3 70B Instruct Meta AI • Dense Transformer | 52.8% | Zero-shot Chain-of-Thought | Dec 2024 | Meta AI Model Card |
| #11 | Claude 3.5 Haiku Anthropic • Dense Transformer | 41.6% | Zero-shot Chain-of-Thought | Nov 2024 | Anthropic Model Card |
| #12 | GPT-4o mini OpenAI • Multimodal Small Transformer | 40.2% | Zero-shot Chain-of-Thought | Jul 2024 | OpenAI Model Card |
Evaluation Methodology & Harness
Diamond subset comprises the highest-confidence, peer-validated 198 questions. Evaluated zero-shot with Chain-of-Thought (CoT).
198 rigorously vetted graduate-level scientific problems requiring multi-step domain deduction.
Why This Benchmark Matters
Measures deep scientific reasoning and resilience against surface-level web search shortcuts.
Known Limitations & Contamination Risks
Small test set (198 questions) means statistical margins of error require careful reporting.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.