DeepsecBench
Code · 2026-07-27
DeepsecBench is Vercel's AI Gateway leaderboard for the open-source deepsec security harness, which points coding agents (the Claude Agent SDK, the Codex SDK, or pi, depending on the model) at a fixed, undisclosed open-source codebase snapshot to find vulnerabilities across 50 entry-point files. Agent findings are matched against a golden set of 231 human-judged issues, additional findings are classified by a judge model, and the published score is 100*5PR/(4P+R), an F2-style combination weighting recall twice as heavily as precision. Each leaderboard entry reports the median of the latest three successful runs for one model and effective reasoning level.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 35.6 |
| Claude Opus 5 | 32.6 |
| GPT-5.6 Luna | 26.8 |
| GPT-5.5 | 21.2 |
| GPT-5.6 Terra | 19.2 |
| Kimi K3 | 17.6 |
| DeepSeek-V4-Flash-0731 | 16.5 |
| Grok 4.5 | 16.5 |
| Gemini 3.6 Flash | 11.6 |
| GLM-5.2 | 10.5 |
| Opus 4.8 | 6.9 |
| Inkling | 6.3 |
| Haiku 4.5 | 4.7 |
| Gemini 3.5 Flash-Lite | 2.1 |