Atlas

Benchmarks

← All benchmarks

Terminal-Bench-Science 0.1

Agents

Terminal-Bench-Science 0.1.0 comprises 70 expert-reviewed research workflows across life, physical, mathematical, engineering and earth sciences. Artificial Analysis runs the complete release with mini-swe-agent and reports pass@1 averaged over three repeats per task; a task passes only if all of its tests pass.

Top models (higher is better)

ModelScore
GPT-6 Astra63.3
Claude Opus 5.561.9
Claude Fable 5.143.3
GPT-6 Sol30.0
Claude Opus 528.6
GPT-5.6 Sol22.4
GPT-5.6 Terra13.8
Qwen3.8 Max (0902)11.9
Muse Spark 1.311.0
Gemini 3.8 Flash10.0
GLM-5.39.5
DeepSeek V4.1 Flash9.0
GPT-6 Luna8.6
Grok 4.66.2
DeepSeek V4 Pro 08135.7
MiMo-V2.6-Pro5.7
GLM-5.3 Flash4.8
GPT-5.6 Luna3.3
Step 5 Preview2.4
MiniMax M30.5
Qwen3.8 27B0.5
Gemini 3.5 Flash-Lite0.0
Inkling0.0
Inkling-Small0.0
Loading Atlas data…