Atlas

Benchmarks

← All benchmarks

CyberBench PoC

Agents · 2026-06-28

CyberBench PoC measures whether an agent can produce raw proof-of-concept input bytes that crash a vulnerable OSS-Fuzz target. Agents receive the project source tree, fuzz target binary, and task metadata, but not the crash input or bug description.

Top models (higher is better)

ModelScore
GPT-5.6 Sol93.2
GPT-5.6 Luna88.1
GPT-5.579.7
Kimi K2.677.3
Kimi K375.0
MiniMax M374.6
GLM-5.273.7
Gemini 3.5 Flash-Lite66.7
GPT-5.464.4
DeepSeek-V4-Flash61.0
DeepSeek-V4-Pro57.6
Gemini 3.5 Flash57.6
Inkling55.6
Grok 4.330.5
Opus 4.726.3
Opus 4.820.3
Gemini 3.6 Flash12.3
Gemini 3.1 Pro Preview3.4
Claude Opus 50.0
Qwen3.7-Plus0.0
Loading Atlas data…