ObviousBench
General QA · 2026-07-16
ObviousBench measures whether language models avoid simple, user-visible mistakes in literal counting, spelling transforms, ordering, negation, formatting, arithmetic, word counting, and basic constraint awareness. Its headline answer pass^3 metric is the percentage of 144 private held-out items for which all three sampled answers are correct.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.1 Pro Preview | 100.0 |
| Gemini 3.5 Flash | 100.0 |
| Gemma 4 31B IT | 100.0 |
| GPT-5 | 100.0 |
| GPT-5.5 | 100.0 |
| o3 | 100.0 |
| Qwen3.5-27B | 100.0 |
| Hy3 | 100.0 |
| Grok 4.5 | 100.0 |
| Claude Fable 5 | 99.3 |
| Opus 4.5 | 99.3 |
| Opus 4.8 | 99.3 |
| Gemini 3 Flash Preview | 99.3 |
| Qwen3.5 397B A17B | 99.3 |
| Grok Build 0.1 | 99.3 |
| Muse Spark 1.1 | 98.6 |
| GPT-5.2 | 98.6 |
| GPT-5.4 | 98.6 |
| O1 | 98.6 |
| O4 Mini | 98.6 |
| Qwen3.5 Plus (2026-02-15) | 98.6 |
| Grok 4.20 | 98.6 |
| Grok 4.3 | 98.6 |
| Opus 4.6 | 97.9 |
| Gemini 3.1 Flash-Lite | 97.9 |
| o3-mini | 97.9 |
| Qwen3.6-Max-Preview | 97.9 |
| Inkling | 97.9 |
| GLM-5.1 | 97.9 |
| Sonnet 4.5 | 97.2 |
| Gemini 2.5 Pro | 97.2 |
| Kimi K2 Thinking | 97.2 |
| Kimi K2.5 | 97.2 |
| GPT-5 Nano | 97.2 |
| Qwen3.5 122B A10B | 97.2 |
| Qwen3.7-Max | 97.2 |
| Sonnet 4.6 | 96.5 |
| Sonnet 5 | 96.5 |
| Gemma 4 26B A4B IT | 96.5 |
| Nemotron 3 Ultra 550B A55B | 96.5 |