Terminal-Bench-Science 0.1
Agents
Terminal-Bench-Science 0.1.0 comprises 70 expert-reviewed research workflows across life, physical, mathematical, engineering and earth sciences. Artificial Analysis runs the complete release with mini-swe-agent and reports pass@1 averaged over three repeats per task; a task passes only if all of its tests pass.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 63.3 |
| Claude Opus 5.5 | 61.9 |
| Claude Fable 5.1 | 43.3 |
| GPT-6 Sol | 30.0 |
| Claude Opus 5 | 28.6 |
| GPT-5.6 Sol | 22.4 |
| GPT-5.6 Terra | 13.8 |
| Qwen3.8 Max (0902) | 11.9 |
| Muse Spark 1.3 | 11.0 |
| Gemini 3.8 Flash | 10.0 |
| GLM-5.3 | 9.5 |
| DeepSeek V4.1 Flash | 9.0 |
| GPT-6 Luna | 8.6 |
| Grok 4.6 | 6.2 |
| DeepSeek V4 Pro 0813 | 5.7 |
| MiMo-V2.6-Pro | 5.7 |
| GLM-5.3 Flash | 4.8 |
| GPT-5.6 Luna | 3.3 |
| Step 5 Preview | 2.4 |
| MiniMax M3 | 0.5 |
| Qwen3.8 27B | 0.5 |
| Gemini 3.5 Flash-Lite | 0.0 |
| Inkling | 0.0 |
| Inkling-Small | 0.0 |