Atlas

Benchmarks

← All benchmarks

SM Bench — Overfit

Safety · 2026-02-01

SM Bench Overfit tests whether surface-level risk cues and policy-shaped phrasing override benign intent, common-sense reasoning, or useful task completion. Scores use difficulty-weighted pass, partial, and fail judgments across 100 fixed prompts; this category receives 2× weight in the overall SM Bench score.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash98.4
Gemini 3.1 Flash-Lite Preview97.8
Opus 4.695.6
Gemini 3 Flash Preview94.0
Grok 4.592.3
Opus 4.791.8
Gemma 4 31B IT91.8
Sonnet 4.691.5
MiMo-V2-Pro90.7
Opus 4.588.5
Opus 4.887.4
MiMo-V2.5-Pro85.8
Gemini 3.1 Pro Preview84.2
Trinity Large Preview83.9
Sonnet 4.583.1
Gemini 3 Pro Preview83.1
Sonnet 583.1
GPT-4o83.1
Haiku 4.582.0
MiMo-V2-Omni81.4
Grok 4.380.9
GLM-5.280.3
MiniMax M379.2
GPT-5.5 Instant79.0
GLM-5.177.6
DeepSeek-V4-Pro76.5
GPT-4.176.5
MiMo-V2.576.0
DeepSeek-R1-052874.9
GLM-4.773.8
Qwen3.7-Max72.7
GLM 5 Turbo70.5
Qwen3.6 Plus Preview69.4
Kimi K2.568.8
Kimi K2.667.2
MiniMax M2.166.7
Qwen Max (source-unspecified mainline)66.7
Grok 4.1 Fast66.1
GPT-4o Mini65.0
GPT-5.6 Sol63.7
Loading Atlas data…