Atlas

Benchmarks

← All benchmarks

CyberBench v1.1 Overall

Agents

CyberBench v1.1 averages the 60-task proof-of-concept and 56-task patch success rates. Vals reran the benchmark using offline task sandboxes and a hardened clean-room grader; these scores are explicitly incomparable with v1. The unequal track sizes mean the aggregate has no single binomial denominator.

Top models (higher is better)

ModelScore
MiMo-V2.6-Flash75.4
DeepSeek V4.1 Flash73.7
MiMo-V2.6-Pro72.9
Muse Spark 1.372.7
Claude Fable 5.170.4
Grok 4.666.0
Claude Opus 565.4
Gemini 3.8 Flash43.8
GPT-6 Astra41.1
Loading Atlas data…