CyberBench v1.1 Overall
Agents
CyberBench v1.1 averages the 60-task proof-of-concept and 56-task patch success rates. Vals reran the benchmark using offline task sandboxes and a hardened clean-room grader; these scores are explicitly incomparable with v1. The unequal track sizes mean the aggregate has no single binomial denominator.
Top models (higher is better)
| Model | Score |
|---|---|
| MiMo-V2.6-Flash | 75.4 |
| DeepSeek V4.1 Flash | 73.7 |
| MiMo-V2.6-Pro | 72.9 |
| Muse Spark 1.3 | 72.7 |
| Claude Fable 5.1 | 70.4 |
| Grok 4.6 | 66.0 |
| Claude Opus 5 | 65.4 |
| Gemini 3.8 Flash | 43.8 |
| GPT-6 Astra | 41.1 |