Atlas

Benchmarks

← All benchmarks

Terminal-Bench 2.1 - Easy

Agents · 2026-05-06

The Easy split of Terminal-Bench 2.1 reports the percentage of easy terminal-use tasks completed successfully; higher scores are better.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Opus 4.8100.0
Claude Opus 5100.0
Sonnet 4.6100.0
Sonnet 5100.0
Composer 2.5100.0
Gemini 3.5 Flash100.0
Gemini 3.6 Flash100.0
GLM-5.2100.0
MiMo-V2.5100.0
Opus 4.791.7
Gemini 3.1 Pro Preview91.7
GPT-5.6 Sol91.7
Grok 4.591.7
Inkling91.7
Kimi K2.7 Code91.7
Kimi K391.7
Nemotron 3 Ultra 550B A55B91.7
Haiku 4.583.3
GLM-5.183.3
GPT-5.4 Nano83.3
GPT-5.583.3
GPT-5.6 Luna83.3
GPT-5.6 Terra83.3
Grok 4.2083.3
Kimi K2.683.3
MiniMax M383.3
Qwen3.7-Max83.3
Gemini 3 Flash Preview75.0
GPT-5.4 Mini75.0
Grok 4.375.0
Laguna M.175.0
MiMo-V2.5-Pro75.0
MiniMax M2.775.0
Muse Spark 1.175.0
Qwen3.6 Plus (2026-04-02)75.0
Qwen3.7-Plus75.0
DeepSeek-V4-Pro66.7
Gemini 3.1 Flash-Lite Preview66.7
Gemini 3.5 Flash-Lite66.7
Loading Atlas data…