Atlas

Benchmarks

← All benchmarks

SM Bench — Ambiguous Interpretation

Safety · 2026-02-01

SM Bench Ambiguous Interpretation tests whether a model asks for clarification when context is genuinely underspecified while avoiding needless refusal or an unwarranted harmful interpretation. Scores are difficulty-weighted judge credit across 100 fixed prompts.

Top models (higher is better)

ModelScore
Claude Fable 592.6
Grok 4.592.0
Sonnet 592.0
GPT-5 Mini91.7
GLM-5.190.8
Gemma 4 31B IT90.8
Opus 4.890.5
GPT-5.190.2
Kimi K2.689.9
Opus 4.589.6
Opus 4.789.3
Kimi K2.589.3
GLM-589.3
Qwen3.6 Plus Preview89.3
Sonnet 4.689.0
GPT-5.289.0
GPT-5.3 Instant88.7
GPT-5.5 Instant88.1
GPT-4.187.8
GPT-5.487.8
MiniMax M387.5
GLM-5.287.2
GPT-5.3-Codex86.9
MiMo-V2.5-Pro86.6
Haiku 4.586.6
GPT-5.6 Terra Pro86.3
GPT-5.4 Mini86.3
GPT-5.4 Nano86.3
Grok 4.386.0
Gemini 3.5 Flash86.0
GPT-5.6 Sol86.0
MiniMax M2.786.0
Gemini 3.1 Pro Preview85.7
MiMo-V2-Pro85.1
Gemini 3 Pro Preview84.5
Opus 4.683.6
GPT-5.6 Luna83.3
Sonnet 4.582.7
MiMo-V2-Omni82.7
GPT-5.582.7
Loading Atlas data…