Atlas

Benchmarks

← All benchmarks

Terminal-Bench 2.1

Agents · 2026-05-06

Terminal-Bench 2.1 is a revision of Terminal-Bench 2.0 that fixes issues in 28 of its 89 command-line tasks and introduces continuous validation. Scores are reported as average task accuracy.

Top models (higher is better)

ModelScore
Claude Mythos Preview92.1
Claude Opus 589.1
GPT-5.6 Sol88.8
Kimi K388.3
Claude Mythos 588.0
Macaron-V1-Venti87.6
GPT-5.6 Terra87.4
GPT-5.6 Luna84.7
Claude Fable 584.3
GPT-5.583.4
Grok 4.583.3
Opus 4.783.1
Opus 4.882.7
GLM-5.282.7
Sonnet 580.5
Muse Spark 1.180.0
DeepSeek-V4-Flash-073178.7
Gemini 3.5 Flash78.7
GPT-5.478.3
Gemini 3.6 Flash78.0
Gemini 3 Pro Preview73.9
Sonnet 4.671.2
Gemini 3.1 Pro Preview70.8
Motif-3-Beta70.8
Laguna S 2.170.2
KAT-Coder-Pro V270.0
Nex-N2-Pro67.8
Kimi K2.7 Code67.4
Agnes 2.5 Pro Alpha67.0
MiniMax M366.0
Kimi K2.665.9
MiMo-V2.5-Pro65.2
DeepSeek-V4-Pro64.8
Inkling-Small64.7
Inkling63.8
MiMo-V2.563.7
Muse Spark62.2
GLM-5.161.8
DeepSeek-V4-Flash61.8
MiMo-V2-Flash61.8
Loading Atlas data…