Atlas

Benchmarks

← All benchmarks

CyberBench v1.1 Patch

Agents

The 56-task patch track of CyberBench v1.1. A patch must compile, remove the original sanitizer failure and match maintainer-fixed reference behavior on hidden holdouts. This revised roster and clean-room grading protocol are distinct from the earlier 59-scored-item track.

Top models (higher is better)

ModelScore
Claude Fable 5.187.5
Gemini 3.8 Flash87.5
Claude Opus 585.7
DeepSeek V4.1 Flash85.7
MiMo-V2.6-Flash85.7
MiMo-V2.6-Pro85.7
GPT-6 Astra82.1
Muse Spark 1.382.1
Grok 4.680.4
Loading Atlas data…