Atlas

Benchmarks

← All benchmarks

CRED v0.92 - Data integrity

Science · 2026-07-20

The CRED v0.92 data-integrity chapter measures detection recall on errors in how source data are selected, transformed, merged, or represented in an empirical policy evaluation.

Top models (higher is better)

ModelScore
Opus 4.1100.0
Sonnet 4.6100.0
GPT-5.5100.0
DeepSeek-V4-Flash100.0
Gemini 3.5 Flash100.0
Opus 4.8100.0
Claude Fable 5100.0
GLM-5.2100.0
Sonnet 5100.0
GPT-5.6 Luna100.0
GPT-5.6 Sol100.0
GPT-5.6 Terra100.0
Kimi K3100.0
Opus 497.5
Sonnet 4.597.5
Opus 4.597.5
Kimi K2.597.5
GLM-597.5
DeepSeek-V4-Pro97.5
Kimi K2 Thinking95.0
Kimi K2.695.0
Ring 2.6 1T92.5
MiniMax M392.5
Gemma 4 31B IT85.0
Hy3 preview85.0
Nemotron 3 Ultra 550B A55B85.0
Qwen3.5 397B A17B82.5
DeepSeek-R1-052880.0
gpt-oss-120b80.0
DeepSeek-V3.280.0
Kimi K2 Instruct75.0
GLM-4.575.0
GLM 4.675.0
GLM-4.775.0
Nemotron 3 Super 120B A12B75.0
DeepSeek-R172.5
DeepSeek-V3.167.5
MiniMax M265.0
DeepSeek-V3-032452.5
O150.0
Loading Atlas data…