Tool & Agentic Evaluation Benchmark

BFCL (Berkeley Function Calling)

Berkeley Function Calling Leaderboard

Comprehensive evaluation of LLM tool-calling and function-calling capabilities across simple, multi-turn, parallel, and executable API scenarios.

Verified Model Leaderboard

Ranked by verified % AST Accuracy score from official technical reports.

Primary Metric: % AST Accuracy
RankAI Model & ProviderScore (% AST Accuracy)Evaluation MethodologyVerification DateSource
#1OpenAI o3-mini

OpenAI Reinforcement Learning Reasoning Model

90.1%Tool-calling evaluation with reasoning effort enabledJan 2025Berkeley Function Calling Leaderboard
#2Claude 3.5 Sonnet (v2)

Anthropic Dense Transformer

89.2%AST accuracy across single and multi-turn tool callingOct 2024Berkeley Function Calling Leaderboard
#3DeepSeek-V3

DeepSeek MoE (Multi-Head Latent Attention)

88.9%AST accuracy on tool calling tasksDec 2024Berkeley Function Calling Leaderboard
#4GPT-4o

OpenAI Multimodal Transformer

88.5%AST accuracy on single & multi-turn tool callingAug 2024Berkeley Function Calling Leaderboard
#5Gemini 2.0 Flash

Google DeepMind Multimodal MoE

87.4%Function calling evaluationDec 2024Berkeley Function Calling Leaderboard
#6Llama 3.3 70B Instruct

Meta AI Dense Transformer

85.6%Tool calling evaluationDec 2024Berkeley Function Calling Leaderboard

Evaluation Methodology & Harness

Measures Abstract Syntax Tree (AST) correctness, parameter extraction, type safety, hallucination avoidance, and live REST API execution.

Dataset & Test Set Construction

Over 2,000 test cases including single-call, multi-turn conversational tool selection, and live REST execution across Java, Python, and REST.

Why This Benchmark Matters

Crucial for autonomous AI agents, enterprise workflow automation, and tool-augmented applications.

Known Limitations & Contamination Risks

Different programming languages and custom schema formatting require standardized tool definitions.

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.