Atlas

Benchmarks

← All benchmarks

ExploitBench (UK AISI/CAISI July preliminary)

Code · 2026-07-23

Normalized exploit-development score on 41 V8 vulnerabilities in the July 23 joint Kimi K3 assessment. The public report does not specify the attempt aggregation or full agent settings; kept separate from CAISI's earlier AutoNudge table and September best-of-three protocol.

Top models (higher is better)

ModelScore
Kimi K332.0
GLM-5.224.0
Loading Atlas data…