Atlas

Benchmarks

← All benchmarks

CRED v0.92

Science · 2026-07-20

CRED v0.92 evaluates automated verifiers on locating known errors injected into 200 copies of 100 self-reproducing empirical policy papers. Each configuration reviews the manuscript, rendered tables, analysis code, bibliography, and reproduction diagnosis; deterministic string matching credits findings that quote the changed content. The headline score is pooled single-error detection recall, with failed calls and refusals counted as misses.

Top models (higher is better)

ModelScore
GPT-5.6 Sol99.0
Kimi K398.0
GPT-5.6 Terra97.0
GPT-5.596.5
Claude Fable 596.5
GPT-5.6 Luna95.0
GLM-5.289.5
Kimi K2.685.5
Opus 4.885.0
Sonnet 4.683.5
DeepSeek-V4-Pro83.5
DeepSeek-V4-Flash80.5
Opus 4.178.0
Sonnet 577.5
Sonnet 4.577.0
Opus 476.5
Opus 4.573.0
GLM-571.5
Nemotron 3 Ultra 550B A55B71.5
MiniMax M371.0
Kimi K2.569.0
Gemini 3.5 Flash68.5
Ring 2.6 1T67.5
Kimi K2 Thinking65.0
GLM-4.761.0
Kimi K2 Instruct58.0
Qwen3.5 397B A17B55.0
DeepSeek-R1-052853.0
Hy3 preview50.5
GLM 4.649.0
Nemotron 3 Super 120B A12B46.5
GLM-4.545.5
gpt-oss-120b45.0
DeepSeek-V3.241.0
DeepSeek-R139.0
Gemma 4 31B IT36.5
DeepSeek-V3.134.5
Mistral Large 3 675B Instruct 251232.0
MiniMax M231.0
O130.0
Loading Atlas data…