SRE Bench — Capability Score
Code · 2026-09-21
Mean fraction of six verified subtasks completed per SRE Bench instance, across 262 instances (1,572 subtasks). Incomplete instances contribute their actual completed subtasks. The evaluation unit is the instance because its six subtasks are dependent.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 63.9 |
| GPT-5.6 Sol | 59.5 |
| Claude Opus 5.5 | 53.6 |
| Claude Fable 5.1 | 36.6 |
| Claude Opus 5 | 31.1 |
| GPT-5.5 | 17.0 |
| DeepSeek V4.1 Flash | 14.2 |
| Gemini 3.7 Flash | 14.2 |
| Hy4 Preview | 10.8 |
| Grok 4.5 | 6.3 |
| GLM-5.2 | 3.4 |