Atlas

Benchmarks

← All benchmarks

SRE Bench — Capability Score

Code · 2026-09-21

Mean fraction of six verified subtasks completed per SRE Bench instance, across 262 instances (1,572 subtasks). Incomplete instances contribute their actual completed subtasks. The evaluation unit is the instance because its six subtasks are dependent.

Top models (higher is better)

ModelScore
GPT-6 Astra63.9
GPT-5.6 Sol59.5
Claude Opus 5.553.6
Claude Fable 5.136.6
Claude Opus 531.1
GPT-5.517.0
DeepSeek V4.1 Flash14.2
Gemini 3.7 Flash14.2
Hy4 Preview10.8
Grok 4.56.3
GLM-5.23.4
Loading Atlas data…