Atlas

Benchmarks

← All benchmarks

CRED v0.92 - Paper–code

Science · 2026-07-20

The CRED v0.92 paper–code inconsistency chapter measures detection recall on errors that become visible only by comparing claims in an empirical policy paper with its analysis code.

Top models (higher is better)

ModelScore
GPT-5.6 Sol98.9
GPT-5.6 Terra95.6
Kimi K395.6
GPT-5.594.4
GPT-5.6 Luna93.3
Claude Fable 592.2
GLM-5.280.0
Kimi K2.675.6
DeepSeek-V4-Pro71.1
Opus 4.871.1
DeepSeek-V4-Flash67.8
Sonnet 4.664.4
Nemotron 3 Ultra 550B A55B56.7
Opus 455.6
Opus 4.155.6
Sonnet 4.555.6
Sonnet 555.6
MiniMax M354.4
Opus 4.550.0
GLM-546.7
Kimi K2 Thinking44.4
Kimi K2.544.4
Kimi K2 Instruct41.1
Ring 2.6 1T41.1
GLM-4.736.7
Gemini 3.5 Flash36.7
DeepSeek-R1-052832.2
Qwen3.5 397B A17B31.1
Hy3 preview26.7
GLM-4.525.6
GLM 4.625.6
gpt-oss-120b22.2
DeepSeek-V3.222.2
Nemotron 3 Super 120B A12B20.0
Mistral Large 3 675B Instruct 251218.9
DeepSeek-R115.6
DeepSeek-V3.114.4
Gemma 3 27B IT10.0
MiniMax M27.8
DeepSeek-V3-03245.6
Loading Atlas data…