Atlas

Benchmarks

← All benchmarks

Terminal-Bench Science 0.1 (Vals)

Science · 2026-09-23

Seventy expert-authored scientific terminal tasks evaluated with Terminus 2 and an eight-hour agent budget. One graded result per task; infrastructure/provider errors are retried, including some successful reattempts, while graded failures are not. This retry-conditioned protocol is separate from the official native-agent three-trial evaluation.

Top models (higher is better)

ModelScore
GPT-6 Astra65.7
Claude Opus 5.548.6
Claude Fable 5.134.3
Claude Opus 522.9
Muse Spark 1.314.3
Claude Fable 512.9
GPT-5.6 Sol12.9
Grok 4.711.4
Opus 4.810.0
GPT-5.6 Terra10.0
Gemini 3.8 Flash8.6
Gemini 3.7 Flash5.7
GPT-5.6 Luna5.7
MiMo-V2.6-Flash5.7
DeepSeek V4.1 Flash4.3
Gemini 3.5 Flash4.3
GLM-5.34.3
Grok 4.64.3
Sonnet 52.9
Kimi K32.9
MiMo-V2.6-Pro2.9
Gemini 3.6 Flash1.4
Hy4 Preview1.4
Qwen3.8 27B1.4
DeepSeek-V4-Flash-07310.0
DeepSeek V4 Pro 08130.0
GLM-5.3 Flash0.0
Mercury 2.50.0
Loading Atlas data…