Atlas

Benchmarks

← All benchmarks

CRED v0.92 - Prose–table

Science · 2026-07-20

The CRED v0.92 prose–table inconsistency chapter measures detection recall on errors exposed by comparing narrative claims with the paper's rendered tables.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Kimi K3100.0
Sonnet 4.698.0
GPT-5.6 Sol98.0
GPT-5.596.0
GPT-5.6 Terra96.0
Opus 4.894.0
GLM-5.294.0
Opus 4.192.0
GPT-5.6 Luna92.0
Sonnet 4.590.0
Kimi K2.690.0
Sonnet 590.0
Opus 488.0
DeepSeek-V4-Pro88.0
Gemini 3.5 Flash88.0
Opus 4.584.0
GLM-584.0
Ring 2.6 1T82.0
DeepSeek-V4-Flash80.0
GLM-4.778.0
Kimi K2.578.0
Nemotron 3 Ultra 550B A55B76.0
MiniMax M372.0
Kimi K2 Thinking64.0
Qwen3.5 397B A17B60.0
Kimi K2 Instruct58.0
DeepSeek-R1-052850.0
GLM 4.650.0
Nemotron 3 Super 120B A12B50.0
Hy3 preview46.0
GLM-4.536.0
gpt-oss-120b36.0
O134.0
Gemma 4 31B IT32.0
DeepSeek-R130.0
DeepSeek-V3.220.0
Mistral Small 3.2 24B Instruct 250618.0
DeepSeek-V3.118.0
MiniMax M218.0
Loading Atlas data…