Atlas

Benchmarks

← All benchmarks

Cybench

Agents · 2024-08-15

A cybersecurity agent benchmark measuring autonomous vulnerability discovery and exploitation across sandboxed challenges.

Top models (higher is better)

ModelScore
Claude Mythos Preview100.0
Opus 4.693.0
Opus 4.582.0
GPT-573.5
Sonnet 4.560.0
gpt-oss-120b49.5
Opus 446.9
Grok 443.0
Opus 4.142.0
DeepSeek-V3.140.0
Grok 4.139.0
DeepSeek-R1-052835.5
Sonnet 435.0
Grok 4 Fast30.0
o3-mini22.5
Claude 3.7 Sonnet20.0
Claude 3.5 Sonnet (June 2024)17.5
GPT-4.517.5
DeepSeek-R116.7
GPT-4o (2024-11-20)12.5
Opus 310.0
o1-mini10.0
o1 Preview10.0
Gemini 1.5 Pro7.5
Llama 3.1 405B Instruct7.5
Mixtral 8x22B Instruct7.5
Llama 3 70B Instruct5.0
Loading Atlas data…