Atlas

Benchmarks

← All benchmarks

CRED v0.92 - Execution

Science · 2026-07-20

The CRED v0.92 execution chapter measures detection recall on failures recorded in the deterministic Replication and Build Defects diagnosis supplied to every verifier.

Top models (higher is better)

ModelScore
GPT-4 Turbo100.0
GPT-4o100.0
GPT-4o Mini100.0
Llama 3.3 70B Instruct100.0
O1100.0
DeepSeek-R1100.0
Gemma 3 27B IT100.0
DeepSeek-V3-0324100.0
Llama 4 Maverick Instruct100.0
GPT-4.1 Nano100.0
GPT-4.1100.0
Qwen3-235B-A22B100.0
Opus 4100.0
DeepSeek-R1-0528100.0
Mistral Small 3.2 24B Instruct 2506100.0
Kimi K2 Instruct100.0
Qwen3 235B A22B Instruct 2507100.0
GLM-4.5100.0
gpt-oss-120b100.0
Opus 4.1100.0
DeepSeek-V3.1100.0
Sonnet 4.5100.0
GLM 4.6100.0
MiniMax M2100.0
Kimi K2 Thinking100.0
Opus 4.5100.0
DeepSeek-V3.2100.0
Mistral Large 3 675B Instruct 2512100.0
GLM-4.7100.0
Kimi K2.5100.0
GLM-5100.0
Sonnet 4.6100.0
Nemotron 3 Super 120B A12B100.0
Gemma 4 31B IT100.0
Kimi K2.6100.0
GPT-5.5100.0
Hy3 preview100.0
DeepSeek-V4-Flash100.0
DeepSeek-V4-Pro100.0
Ring 2.6 1T100.0
Loading Atlas data…