Atlas

Benchmarks

← All benchmarks

ObviousBench

General QA · 2026-07-16

ObviousBench measures whether language models avoid simple, user-visible mistakes in literal counting, spelling transforms, ordering, negation, formatting, arithmetic, word counting, and basic constraint awareness. Its headline answer pass^3 metric is the percentage of 144 private held-out items for which all three sampled answers are correct.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview100.0
Gemini 3.5 Flash100.0
Gemma 4 31B IT100.0
GPT-5100.0
GPT-5.5100.0
o3100.0
Qwen3.5-27B100.0
Hy3100.0
Grok 4.5100.0
Claude Fable 599.3
Opus 4.599.3
Opus 4.899.3
Gemini 3 Flash Preview99.3
Qwen3.5 397B A17B99.3
Grok Build 0.199.3
Muse Spark 1.198.6
GPT-5.298.6
GPT-5.498.6
O198.6
O4 Mini98.6
Qwen3.5 Plus (2026-02-15)98.6
Grok 4.2098.6
Grok 4.398.6
Opus 4.697.9
Gemini 3.1 Flash-Lite97.9
o3-mini97.9
Qwen3.6-Max-Preview97.9
Inkling97.9
GLM-5.197.9
Sonnet 4.597.2
Gemini 2.5 Pro97.2
Kimi K2 Thinking97.2
Kimi K2.597.2
GPT-5 Nano97.2
Qwen3.5 122B A10B97.2
Qwen3.7-Max97.2
Sonnet 4.696.5
Sonnet 596.5
Gemma 4 26B A4B IT96.5
Nemotron 3 Ultra 550B A55B96.5
Loading Atlas data…