CRED v0.92 - Prose–table
Science · 2026-07-20
The CRED v0.92 prose–table inconsistency chapter measures detection recall on errors exposed by comparing narrative claims with the paper's rendered tables.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 100.0 |
| Kimi K3 | 100.0 |
| Sonnet 4.6 | 98.0 |
| GPT-5.6 Sol | 98.0 |
| GPT-5.5 | 96.0 |
| GPT-5.6 Terra | 96.0 |
| Opus 4.8 | 94.0 |
| GLM-5.2 | 94.0 |
| Opus 4.1 | 92.0 |
| GPT-5.6 Luna | 92.0 |
| Sonnet 4.5 | 90.0 |
| Kimi K2.6 | 90.0 |
| Sonnet 5 | 90.0 |
| Opus 4 | 88.0 |
| DeepSeek-V4-Pro | 88.0 |
| Gemini 3.5 Flash | 88.0 |
| Opus 4.5 | 84.0 |
| GLM-5 | 84.0 |
| Ring 2.6 1T | 82.0 |
| DeepSeek-V4-Flash | 80.0 |
| GLM-4.7 | 78.0 |
| Kimi K2.5 | 78.0 |
| Nemotron 3 Ultra 550B A55B | 76.0 |
| MiniMax M3 | 72.0 |
| Kimi K2 Thinking | 64.0 |
| Qwen3.5 397B A17B | 60.0 |
| Kimi K2 Instruct | 58.0 |
| DeepSeek-R1-0528 | 50.0 |
| GLM 4.6 | 50.0 |
| Nemotron 3 Super 120B A12B | 50.0 |
| Hy3 preview | 46.0 |
| GLM-4.5 | 36.0 |
| gpt-oss-120b | 36.0 |
| O1 | 34.0 |
| Gemma 4 31B IT | 32.0 |
| DeepSeek-R1 | 30.0 |
| DeepSeek-V3.2 | 20.0 |
| Mistral Small 3.2 24B Instruct 2506 | 18.0 |
| DeepSeek-V3.1 | 18.0 |
| MiniMax M2 | 18.0 |