BFCL (Berkeley Function Calling)
Berkeley Function Calling Leaderboard
Comprehensive evaluation of LLM tool-calling and function-calling capabilities across simple, multi-turn, parallel, and executable API scenarios.
Verified Model Leaderboard
Ranked by verified % AST Accuracy score from official technical reports.
| Rank | AI Model & Provider | Score (% AST Accuracy) | Evaluation Methodology | Verification Date | Source |
|---|---|---|---|---|---|
| #1 | OpenAI o3-mini OpenAI • Reinforcement Learning Reasoning Model | 90.1% | Tool-calling evaluation with reasoning effort enabled | Jan 2025 | Berkeley Function Calling Leaderboard |
| #2 | Claude 3.5 Sonnet (v2) Anthropic • Dense Transformer | 89.2% | AST accuracy across single and multi-turn tool calling | Oct 2024 | Berkeley Function Calling Leaderboard |
| #3 | DeepSeek-V3 DeepSeek • MoE (Multi-Head Latent Attention) | 88.9% | AST accuracy on tool calling tasks | Dec 2024 | Berkeley Function Calling Leaderboard |
| #4 | GPT-4o OpenAI • Multimodal Transformer | 88.5% | AST accuracy on single & multi-turn tool calling | Aug 2024 | Berkeley Function Calling Leaderboard |
| #5 | Gemini 2.0 Flash Google DeepMind • Multimodal MoE | 87.4% | Function calling evaluation | Dec 2024 | Berkeley Function Calling Leaderboard |
| #6 | Llama 3.3 70B Instruct Meta AI • Dense Transformer | 85.6% | Tool calling evaluation | Dec 2024 | Berkeley Function Calling Leaderboard |
Evaluation Methodology & Harness
Measures Abstract Syntax Tree (AST) correctness, parameter extraction, type safety, hallucination avoidance, and live REST API execution.
Over 2,000 test cases including single-call, multi-turn conversational tool selection, and live REST execution across Java, Python, and REST.
Why This Benchmark Matters
Crucial for autonomous AI agents, enterprise workflow automation, and tool-augmented applications.
Known Limitations & Contamination Risks
Different programming languages and custom schema formatting require standardized tool definitions.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.