Atlas

Benchmarks

← All benchmarks

Agentic Tool Use (Enterprise)

Agents · 2024-09-19

Score on ToolComp-Enterprise, which tests dependent multi-tool workflows across 11 tools.

Top models (higher is better)

ModelScore
O170.1
Gemini 2.5 Pro Preview 03-2568.8
O1 Pro67.0
o1 Preview66.4
DeepSeek-R165.3
o3-mini65.3
Claude 3.7 Sonnet65.3
GPT-4o64.6
DeepSeek-V3-032464.2
GPT-4.563.8
Gemini 2.0 Flash Thinking Experimental 01-2163.2
DeepSeek-V362.5
Gemini 2.0 Pro Experimental 02-0561.5
GPT-4 0125 Preview60.8
Gemini 1.5 Pro Experimental 082760.3
Gemini 2.0 Flash Experimental60.1
GPT-4o (2024-08-06)59.9
Claude 3.5 Sonnet (June 2024)59.4
Gemini 2.0 Flash 00155.6
Claude 3 Sonnet54.2
Gemini 2.0 Flash-Lite53.8
Opus 352.8
GPT-4o Mini51.7
GPT-451.4
Llama 3.1 405B Instruct50.4
Mistral Large 2 (Instruct 2407)50.4
Nova Pro47.0
Gemini 1.5 Pro Preview 051440.4
Llama 3.1 70B Instruct37.2
Nova Lite36.5
C4AI Command R+30.2
Nova Micro28.9
Llama 3.1 8B Instruct17.4
Loading Atlas data…