Atlas

Benchmarks

← All benchmarks

DeepsecBench — Revalidated V1

Code

Vercel DeepsecBench with the dub-sol-revalidated-v1 oracle of 232 issues and score-v2-binary grading. Coding agents inspect a fixed codebase; the leaderboard selects the median F2 run from the latest three successful runs per agent, model and effective reasoning level. Score is 100*5PR/(4P+R). This revised oracle is kept separate from the earlier 231-issue release.

Top models (higher is better)

ModelScore
GPT-6 Sol40.9
GPT-6 Astra37.8
GPT-5.6 Sol35.4
Claude Opus 532.4
GPT-5.6 Luna26.7
Claude Opus 5.526.7
GLM-5.321.9
GPT-5.521.1
GPT-6 Luna21.1
GPT-5.6 Terra19.1
Kimi K317.5
Sonnet 517.0
DeepSeek-V4-Flash-073116.5
Qwen3.8-Max16.5
Grok 4.516.5
Grok 4.616.2
Gemini 3.6 Flash11.5
Grok 4.710.9
GLM-5.210.4
Opus 4.86.9
Inkling6.3
Haiku 4.54.7
Gemini 3.5 Flash-Lite2.1
Loading Atlas data…