Atlas

Benchmarks

← All benchmarks

Terminal-Bench 4.0 — Hardware

Agents · 2026-09-22

Resolution rate on the five hardware tasks in Terminal-Bench 4.0.

Top models (higher is better)

ModelScore
GPT-6 Astra73.3
Opus 4.820.0
GPT-5.6 Sol20.0
MiMo-V2.6-Pro20.0
Gemini 3.8 Flash13.3
GLM-5.3 Flash13.3
GLM-5.313.3
Grok 4.713.3
MiMo-V2.6-Flash13.3
Sonnet 56.7
DeepSeek-V4-Flash-07316.7
GPT-5.6 Terra6.7
Grok 4.66.7
DeepSeek V4.1 Flash0.0
DeepSeek V4 Pro 08130.0
Gemini 3.5 Flash0.0
Gemini 3.6 Flash0.0
Gemini 3.7 Flash0.0
GPT-5.6 Luna0.0
Grok 4.50.0
Hy4 Preview0.0
Inkling0.0
Inkling-Small0.0
Kimi K30.0
Mercury 2.50.0
Muse Spark 1.20.0
Muse Spark 1.30.0
Qwen3.8 27B0.0
Loading Atlas data…