Toolathlon
Agents · 2025-10-29
Toolathlon evaluates agents on 108 real-world tasks spanning software, productivity and communication applications reached through their own tools and APIs, with Pass@1 reported as mean task success over three runs. This row holds the original score series, as released in October 2025 and reported through June 2026. Toolathlon-Verified (2026-06-30) revised the prompts, ground truths and evaluators of the same 108 tasks and, in the maintainers' own words, "begins a new official score series, so these results are not directly comparable with scores from earlier versions"; those scores live on the toolathlon-verified row.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.5 Flash | 56.5 |
| GPT-5.5 | 55.6 |
| GPT-5.4 | 54.6 |
| GPT-5.3-Codex | 51.9 |
| Gemini 3 Flash Preview | 49.4 |
| Sonnet 4.5 | 38.9 |
| Gemini 3 Pro Preview | 36.4 |
| Gemini 2.5 Pro | 10.5 |
| Gemini 2.5 Flash | 3.7 |