DeepsecBench — Revalidated V1
Code
Vercel DeepsecBench with the dub-sol-revalidated-v1 oracle of 232 issues and score-v2-binary grading. Coding agents inspect a fixed codebase; the leaderboard selects the median F2 run from the latest three successful runs per agent, model and effective reasoning level. Score is 100*5PR/(4P+R). This revised oracle is kept separate from the earlier 231-issue release.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-6 Sol | 40.9 |
| GPT-6 Astra | 37.8 |
| GPT-5.6 Sol | 35.4 |
| Claude Opus 5 | 32.4 |
| GPT-5.6 Luna | 26.7 |
| Claude Opus 5.5 | 26.7 |
| GLM-5.3 | 21.9 |
| GPT-5.5 | 21.1 |
| GPT-6 Luna | 21.1 |
| GPT-5.6 Terra | 19.1 |
| Kimi K3 | 17.5 |
| Sonnet 5 | 17.0 |
| DeepSeek-V4-Flash-0731 | 16.5 |
| Qwen3.8-Max | 16.5 |
| Grok 4.5 | 16.5 |
| Grok 4.6 | 16.2 |
| Gemini 3.6 Flash | 11.5 |
| Grok 4.7 | 10.9 |
| GLM-5.2 | 10.4 |
| Opus 4.8 | 6.9 |
| Inkling | 6.3 |
| Haiku 4.5 | 4.7 |
| Gemini 3.5 Flash-Lite | 2.1 |