Atlas

Benchmarks

← All benchmarks

Terminal-Bench 4.0 — Operations

Agents · 2026-09-22

Resolution rate on the nine operations tasks in Terminal-Bench 4.0.

Top models (higher is better)

ModelScore
GPT-6 Astra44.4
GLM-5.329.6
Grok 4.722.2
MiMo-V2.6-Pro22.2
GPT-5.6 Sol18.5
GPT-5.6 Terra18.5
Grok 4.618.5
GLM-5.3 Flash14.8
Gemini 3.7 Flash11.1
Gemini 3.8 Flash11.1
MiMo-V2.6-Flash11.1
Muse Spark 1.311.1
Sonnet 57.4
Grok 4.57.4
Opus 4.83.7
DeepSeek-V4-Flash-07313.7
DeepSeek V4 Pro 08133.7
Gemini 3.5 Flash3.7
Gemini 3.6 Flash3.7
Hy4 Preview3.7
DeepSeek V4.1 Flash0.0
GPT-5.6 Luna0.0
Inkling0.0
Inkling-Small0.0
Kimi K30.0
Mercury 2.50.0
Muse Spark 1.20.0
Qwen3.8 27B0.0
Loading Atlas data…