Atlas

Benchmarks

← All benchmarks

SEC-Bench Pro (CAISI 183 tasks)

Code · 2026-09-17

CAISI's 183 V8 and SpiderMonkey vulnerability tasks. The agent receives a vulnerable source tree, files to audit and flaw type, and must produce a crash that triggers the intended bug. ReAct with bash, Python and continuation nudges; 200-turn limit.

Top models (higher is better)

ModelScore
GLM-5.340.4
Loading Atlas data…