Atlas

Benchmarks

← All benchmarks

CyberBench v1.1 PoC

Agents

The 60-task proof-of-concept track of CyberBench v1.1, evaluated in offline sandboxes with a hardened clean-room grader. A submitted input must crash the vulnerable build and, when the differential oracle is valid, avoid crashing the fixed build.

Top models (higher is better)

ModelScore
MiMo-V2.6-Flash65.0
Muse Spark 1.363.3
DeepSeek V4.1 Flash61.7
MiMo-V2.6-Pro60.0
Claude Fable 5.153.3
Grok 4.651.7
Claude Opus 545.0
Gemini 3.8 Flash0.0
GPT-6 Astra0.0
Loading Atlas data…