Atlas

Benchmarks

← All benchmarks

Aikido Known-CVE Detection (Single-Run Average)

Code · 2026-07-16

The single-run companion to Aikido Known-CVE Detection reports mean recall across three independent audits of the same 26 known vulnerabilities. It measures average one-run consistency rather than the pooled union used by the primary benchmark and is retained as a diagnostic metric.

Top models (higher is better)

ModelScore
GPT-5.6 Terra83.3
GPT-5.6 Sol78.2
GPT-5.6 Luna74.4
GPT-5.573.1
Grok 4.566.7
GPT-5.4 Mini64.1
GPT-5.4 Nano57.7
Opus 4.855.1
Gemini 3.1 Pro Preview47.4
GLM-5.246.2
Opus 4.744.9
Gemini 3.5 Flash44.9
Haiku 4.541.0
Loading Atlas data…