Atlas

Benchmarks

← All benchmarks

CyberGym

Agents · 2025-06-03

CyberGym Level 1 measures whether an AI agent can reproduce a target real-world vulnerability by generating a working proof of concept from its description and unpatched codebase. The official leaderboard's default one-trial setting covers 1,507 historical vulnerabilities from 188 open-source projects.

Top models (higher is better)

ModelScore
Opus 4.689.6
GLM-5.184.0
Claude Mythos 583.8
GPT-5.6 Sol83.6
Gemini 3.5 Flash Cyber83.2
Claude Mythos Preview83.1
GPT-5.581.8
GPT-5.479.0
MiniMax M373.1
Sonnet 4.665.2
GPT-560.2
Muse Spark 1.159.0
DeepSeek-V4-Pro57.7
Opus 4.550.6
Muse Spark43.5
GLM-543.2
Kimi K2.541.3
Gemini 3.1 Pro Preview38.8
Sonnet 4.528.9
Opus 4.125.0
GLM-4.723.5
Sonnet 422.6
Claude 3.7 Sonnet14.5
GPT-4.19.4
Gemini 2.5 Flash Preview 04-174.8
DeepSeek-V33.6
O4 Mini2.5
Qwen3-235B-A22B1.9
Loading Atlas data…