Atlas

Benchmarks

← All benchmarks

Terminal-Bench 4.0 (Artificial Analysis)

Agents

Artificial Analysis' implementation of all 66 Terminal-Bench 4.0 tasks, reporting pass@1 averaged over three repeats per task. Version 4.0 changes task membership, compute and time allowances, instructions and verification; it is separate from Terminal-Bench 2.1.

Top models (higher is better)

ModelScore
Claude Opus 5.559.6
GPT-6 Astra59.6
Claude Fable 5.155.1
Claude Opus 549.0
GPT-6 Sol43.9
Claude Fable 542.4
GLM-5.341.9
GPT-5.6 Sol39.9
Qwen3.8 Max (0902)38.9
GPT-5.6 Terra35.4
MiMo-V2.6-Pro34.8
Muse Spark 1.333.3
Step 5 Preview33.3
GLM-5.3 Flash32.8
DeepSeek V4.1 Flash26.8
Grok 4.725.8
Qwen3.8-Flash-Next25.3
Opus 4.821.7
Grok 4.621.2
Gemini 3.8 Flash19.7
Qwen3.8-Max18.7
DeepSeek-V4-Pro14.6
GPT-5.514.6
Sonnet 514.1
DeepSeek V4 Pro 081314.1
Gemini 3.7 Flash13.6
GPT-5.5 Instant (2026-06-25 hosted snapshot)12.6
GPT-6 Luna12.6
Kimi K312.6
DeepSeek-V4-Flash-073112.1
DeepSeek V4 Flash Vision Exp12.1
GPT-5.6 Luna11.6
Qwen3.8 2.4T A95B11.1
Grok 4.510.6
Agnes 3.0 Flash (hosted)7.1
Gemini 3.6 Flash7.1
Muse Spark 1.27.1
Gemini 3.5 Flash6.6
Muse Spark 1.16.1
Qwen3.8 27B5.6
Loading Atlas data…