Atlas

Benchmarks

← All benchmarks

Toolathlon

Agents · 2025-10-29

Toolathlon evaluates agents on 108 real-world tasks spanning software, productivity and communication applications reached through their own tools and APIs, with Pass@1 reported as mean task success over three runs. This row holds the original score series, as released in October 2025 and reported through June 2026. Toolathlon-Verified (2026-06-30) revised the prompts, ground truths and evaluators of the same 108 tasks and, in the maintainers' own words, "begins a new official score series, so these results are not directly comparable with scores from earlier versions"; those scores live on the toolathlon-verified row.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash56.5
GPT-5.555.6
GPT-5.454.6
GPT-5.3-Codex51.9
Gemini 3 Flash Preview49.4
Sonnet 4.538.9
Gemini 3 Pro Preview36.4
Gemini 2.5 Pro10.5
Gemini 2.5 Flash3.7
Loading Atlas data…