Atlas

Benchmarks

← All benchmarks

Terminal-Bench Hard

Agents · 2025-12-22

Terminal-Bench Hard is the difficult subset of a terminal agent benchmark covering software engineering, system administration, data processing, compilation, model training, debugging, and related shell workflows. Higher scores mean more terminal tasks completed.

Top models (higher is better)

ModelScore
GPT-5.6 Sol65.9
Claude Fable 562.9
GPT-5.6 Terra62.9
GPT-5.560.6
Opus 4.858.3
GPT-5.457.6
Opus 4.754.5
Gemini 3.1 Pro Preview53.8
Sonnet 4.653.0
GPT-5.3-Codex53.0
GPT-5.4 Mini52.3
GLM-5.250.8
Qwen3.7-Max50.8
KAT-Coder-Pro V249.2
Opus 4.648.5
Opus 4.547.0
GPT-5.247.0
Qwen3.7-Plus47.0
DeepSeek-V4-Pro46.2
Gemini 3.5 Flash46.2
GPT-5.145.5
Muse Spark45.5
Kimi K2.7 Code44.7
Kimi K2.643.9
Qwen3.6-Max-Preview43.9
Qwen3.6 Plus (2026-04-02)43.9
GLM-5.143.2
GLM-543.2
MiMo-V2.5-Pro43.2
GPT-5.4 Nano42.4
GPT-5.5 Instant42.4
MiniMax M342.4
Gemini 3 Pro Preview41.7
MiMo-V2.541.7
Grok 4.2040.9
Qwen3.5 397B A17B40.9
MiMo-V2-Pro40.9
MiniMax M2.739.4
DeepSeek-V4-Flash38.6
Gemini 3 Flash Preview38.6
Loading Atlas data…