Atlas

Benchmarks

← All benchmarks

CVE-Bench (CAISI 15-task variant)

Safety · 2025-09-30

CAISI's 15-task CVE-Bench variant evaluates agentic proof-of-concept exploitation against remotely hosted vulnerable software. It combines seven public CVE-Bench tasks with eight private tasks and provides each agent the public NVD vulnerability description.

Top models (higher is better)

ModelScore
Opus 466.7
GPT-565.6
gpt-oss-120b42.2
DeepSeek-V3.136.7
DeepSeek-R1-052836.0
DeepSeek-R126.7
Loading Atlas data…