Atlas

Benchmarks

← All benchmarks

Terminal-Bench 2.1

Agents · 2026-05-06

Terminal-Bench 2.1 is a revision of Terminal-Bench 2.0 that fixes issues in 28 of its 89 command-line tasks and introduces continuous validation. Scores are reported as average task accuracy.

Top models (higher is better)

ModelScore
Claude Mythos Preview92.1
Claude Opus 589.1
Kimi K388.3
Claude Mythos 588.0
Macaron-V1-Venti87.6
GPT-5.6 Terra87.4
GPT-5.6 Sol85.8
GPT-5.6 Luna84.7
Claude Fable 584.3
GPT-5.583.4
Grok 4.583.3
Opus 4.783.1
Opus 4.882.7
GLM-5.282.7
Sonnet 580.5
Muse Spark 1.180.0
DeepSeek-V4-Flash-073178.7
Gemini 3.5 Flash78.7
GPT-5.478.3
Gemini 3.6 Flash78.0
Gemini 3 Pro Preview73.9
Sonnet 4.671.2
Gemini 3.1 Pro Preview70.8
Motif-3-Beta70.8
Laguna S 2.170.2
KAT-Coder-Pro V270.0
Nex-N2-Pro67.8
Kimi K2.7 Code67.4
Agnes 2.5 Pro Alpha67.0
MiniMax M366.0
Kimi K2.665.9
MiMo-V2.5-Pro65.2
DeepSeek-V4-Pro64.8
Inkling-Small64.7
Inkling63.8
MiMo-V2.563.7
Muse Spark62.2
GLM-5.161.8
DeepSeek-V4-Flash61.8
MiMo-V2-Flash61.8
Loading Atlas data…