Atlas

Benchmarks

← All benchmarks

Toolathlon-Verified

Agents · 2026-06-30

Toolathlon-Verified evaluates agents on the verified 108-task Toolathlon suite across real-world applications and tools. The official trajectory release reports Pass@1 as mean task success over three runs per model.

Top models (higher is better)

ModelScore
Claude Opus 580.6
Claude Mythos 579.3
Claude Fable 577.9
Kimi K376.5
GPT-5.6 Sol74.9
Sonnet 574.7
GPT-5.573.5
GPT-5.6 Luna67.9
Gemini 3.5 Flash67.3
Gemini 3.1 Pro Preview61.1
GLM-5.259.9
Kimi K2.7 Code58.0
Kimi K2.658.0
Gemini 3.5 Flash-Lite57.1
Inkling-Small54.4
DeepSeek-V4-Flash50.9
Laguna S 2.149.7
MiMo-V2.549.1
MiniMax M2.747.5
Inkling45.5
Qwen3.5 397B A17B40.7
Nemotron 3 Ultra 550B A55B34.3
Kimi K2.533.0
Haiku 4.526.9
Loading Atlas data…