ExploitBench (UK AISI/CAISI July preliminary)
Code · 2026-07-23
Normalized exploit-development score on 41 V8 vulnerabilities in the July 23 joint Kimi K3 assessment. The public report does not specify the attempt aggregation or full agent settings; kept separate from CAISI's earlier AutoNudge table and September best-of-three protocol.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 32.0 |
| GLM-5.2 | 24.0 |