Atlas
/
Benchmarks
137 sources · 855 models × 1278 benchmarks · 74,806 scores
☾
☀
Capabilities
Models
Compare
Benchmarks
Coverage
Methodology
← All benchmarks
BFCL v4 Overall Accuracy
Agents · 2025-07-17
Overall accuracy on the BFCL v4 function-calling benchmark.
Top models
(higher is better)
Model
Score
Opus 4.5
77.5
Sonnet 4.5
73.2
Gemini 3 Pro Preview
72.5
GLM 4.6
72.4
Grok 4.1 Fast
69.6
Haiku 4.5
68.7
O3 (2025-04-16)
63.0
Grok 4
63.0
Kimi K2 Instruct
59.1
DeepSeek-V3.2-Exp
56.7
Gemini 2.5 Flash
56.2
GPT-5.2 (2025-12-11)
55.9
GPT-5 Mini (2025-08-07)
55.5
GPT-4.1 (2025-04-14)
54.0
O4 Mini (2025-04-16)
53.2
Qwen3 235B A22B Instruct 2507
52.1
GPT-5 Nano (2025-08-07)
51.5
Nanbeige4-3B-Thinking-2511
51.4
GPT-4.1 Mini (2025-04-14)
50.5
Qwen3 32B
48.7
Command A
46.5
Qwen3 8B
42.6
Qwen3 30B A3B Instruct 2507
41.4
Qwen3 14B
41.0
Mistral Large 2.1 (Instruct 2411)
38.4
Mistral Medium 3
37.7
Llama 4 Maverick Instruct FP8
37.3
Mistral Small 3.2 24B Instruct 2506
37.1
Gemini 2.5 Flash-Lite
36.9
Qwen3 4B Instruct 2507
35.7
GPT-4.1 Nano (2025-04-14)
33.0
Command R7B (Dec 2024)
32.1
Llama 3.3 70B Instruct
31.9
Gemma 3 12B IT
30.4
Gemma 3 27B IT
29.5
Phi-4
28.8
Qwen3 1.7B
28.4
Llama 4 Scout Instruct
28.1
Mistral NeMo Instruct 2407
27.6
Nova 2 Lite
27.1
Loading Atlas data…