Atlas

Benchmarks

← All benchmarks

CyberBench Patch

Agents · 2026-06-28

CyberBench Patch measures whether an agent can produce a source-code fix for an OSS-Fuzz crash regression. A successful patch blocks the vulnerability while preserving benign behavior under Vals AI's task checks.

Top models (higher is better)

ModelScore
Gemini 3.6 Flash84.7
GPT-5.6 Sol83.1
Kimi K383.1
Opus 4.881.4
Claude Opus 581.4
GPT-5.481.4
GPT-5.581.4
Gemini 3.5 Flash81.0
GLM-5.281.0
DeepSeek-V4-Pro79.7
GPT-5.6 Luna79.7
Opus 4.778.0
Gemini 3.5 Flash-Lite78.0
MiniMax M378.0
Qwen3.7-Plus75.0
Inkling72.7
DeepSeek-V4-Flash72.4
Kimi K2.671.4
Gemini 3.1 Pro Preview69.5
Grok 4.364.4
Loading Atlas data…