Atlas

Benchmarks

← All benchmarks

Terminal-Bench 4.0

Agents · 2026-08-28

Terminal-Bench 4.0 measures agents completing terminal tasks in Harbor. This major revision removes eight tasks, fixes nineteen tasks, calibrates CPU and memory resources, and gives each task an eight-hour agent timeout. Scores are mean resolution rates across repeated trials; they are not directly comparable to earlier Terminal-Bench task sets.

Top models (higher is better)

ModelScore
GPT-6 Astra58.2
Claude Fable 5.157.9
Claude Opus 553.9
Claude Fable 544.5
GLM-5.341.8
Grok 4.737.6
GPT-5.6 Sol37.3
Muse Spark 1.327.8
GPT-5.6 Terra26.3
MiMo-V2.6-Pro24.7
Opus 4.823.6
MiMo-V2.6-Flash21.2
Grok 4.620.3
GLM-5.3 Flash19.7
Gemini 3.8 Flash19.1
GPT-5.6 Luna17.3
Kimi K312.6
Grok 4.512.4
Sonnet 512.4
DeepSeek V4.1 Flash11.6
Gemini 3.7 Flash11.2
DeepSeek-V4-Flash-07319.1
Muse Spark 1.25.6
Hy4 Preview5.1
Gemini 3.6 Flash4.5
Gemini 3.5 Flash4.0
Qwen3.8 27B4.0
Inkling-Small1.5
DeepSeek V4 Pro 08131.0
Inkling0.0
Mercury 2.50.0
Loading Atlas data…