Atlas

Benchmarks

← All benchmarks

CyberBench Overall

Agents · 2026-06-28

CyberBench Overall is Vals AI's aggregate accuracy across the PoC and Patch tracks. The June 2026 preview evaluates agents with Mini-SWE-Agent on 60 OSS-Fuzz crash-regression tasks harvested after CyberGym's original task window.

Top models (higher is better)

ModelScore
GPT-5.6 Sol88.1
GPT-5.6 Luna83.9
GPT-5.580.5
Kimi K379.0
GLM-5.277.4
MiniMax M376.3
Kimi K2.674.4
GPT-5.472.9
Gemini 3.5 Flash-Lite72.3
Gemini 3.5 Flash69.3
DeepSeek-V4-Pro68.6
DeepSeek-V4-Flash66.7
Inkling64.1
Opus 4.752.1
Opus 4.850.8
Gemini 3.6 Flash48.5
Grok 4.347.5
Claude Opus 540.7
Qwen3.7-Plus37.5
Gemini 3.1 Pro Preview36.4
Loading Atlas data…