Atlas

Benchmarks

← All benchmarks

SM Bench — Anti-hallucination

Safety · 2026-02-01

SM Bench Anti-hallucination tests whether models reject invented facts, acronyms, and unsupported premises instead of responding with fabricated certainty. Scores are difficulty-weighted judge credit across 100 fixed prompts.

Top models (higher is better)

ModelScore
MiMo-V2-Pro100.0
Opus 4.6100.0
Qwen3.7-Max100.0
Opus 4.8100.0
Sonnet 4.6100.0
Sonnet 5100.0
Gemini 3.1 Pro Preview99.0
Opus 4.599.0
Sonnet 4.599.0
MiMo-V2.5-Pro98.4
Claude Fable 598.4
Qwen3.5 397B A17B98.4
Grok 4.597.9
Grok 4.397.9
Opus 4.797.4
Gemini 3 Pro Preview97.4
Grok 4.1 Fast96.9
Gemma 4 31B IT96.9
Qwen3.6 Plus Preview96.9
GLM-5.296.3
Kimi K2.696.3
MiniMax M396.3
GPT-5.5 Instant95.8
Qwen3.5 Plus (2026-02-15)95.3
Gemini 3.1 Flash-Lite Preview94.8
GPT-5.594.8
MiMo-V2-Omni94.2
DeepSeek-V4-Flash94.2
GPT-5.6 Sol Pro94.2
GPT-5.3-Codex93.7
Kimi K2.593.2
GLM-592.7
Haiku 4.592.2
GPT-5.191.6
GPT-4.191.1
DeepSeek-V3.291.1
GPT-5.6 Terra91.1
GPT-5.3 Instant90.6
GPT-5.6 Sol90.6
GPT-5.6 Terra Pro90.6
Loading Atlas data…