Atlas

Benchmarks

← All benchmarks

Agentic Tool Use (Chat)

Agents · 2024-09-19

Score on ToolComp-Chat, which tests dependent tool use with Python and Google Search.

Top models (higher is better)

ModelScore
o3-mini63.5
Gemini 2.5 Pro Preview 03-2562.4
O1 Pro61.4
DeepSeek-R160.9
O160.4
DeepSeek-V358.5
Gemini 2.0 Pro Experimental 02-0557.9
Gemini 2.0 Flash Thinking Experimental 01-2157.4
GPT-4o (2024-08-06)56.9
GPT-4.556.3
Claude 3.7 Sonnet56.3
Claude 3.5 Sonnet (June 2024)56.1
o1 Preview55.1
DeepSeek-V3-032454.3
Gemini 2.0 Flash Experimental53.3
GPT-4 0125 Preview53.0
Gemini 1.5 Pro Experimental 082751.3
Gemini 2.0 Flash-Lite50.3
GPT-4o49.5
Nova Pro49.2
Opus 348.5
Gemini 2.0 Flash 00148.2
Nova Lite41.7
Nova Micro41.0
Claude 3 Sonnet40.4
Mistral Large 2 (Instruct 2407)40.4
Llama 3.1 405B Instruct40.1
GPT-437.9
Gemini 1.5 Pro (May 2024)35.5
Llama 3.1 70B Instruct33.5
GPT-4o Mini32.8
C4AI Command R+20.2
Llama 3.1 8B Instruct6.1
Loading Atlas data…