Atlas

Benchmarks

← All benchmarks

SM Bench

Safety · 2026-02-01

Safetymaxxed Bench (SM Bench) evaluates safety overfitting and useful alignment on a fixed v1 battery of 800 prompts across eight categories. Gemini 3 Flash Preview judges each response as pass, partial, or fail; item credit is weighted 1×, 2×, or 3× for easy, medium, or hard cases, and the overall score is the category-weighted mean with Overfit weighted 2×. Higher is better.

Top models (higher is better)

ModelScore
Grok 4.592.1
Grok 4.388.9
Gemini 3 Flash Preview85.3
Grok 4.1 Fast84.8
GLM-5.184.1
Opus 4.784.0
Gemini 3.1 Flash-Lite Preview83.8
Gemini 3.1 Pro Preview83.4
Kimi K2.582.3
Gemma 4 31B IT82.2
Gemini 3.5 Flash82.0
MiMo-V2-Pro81.5
DeepSeek-V4-Pro81.5
Opus 4.581.4
GLM-5.281.3
Opus 4.680.8
Sonnet 4.579.8
Gemini 3 Pro Preview79.3
MiMo-V2-Omni79.2
Kimi K2.678.5
GPT-4.177.3
GLM-577.0
Qwen3.7-Max76.8
Qwen3.6 Plus Preview74.6
MiMo-V2.5-Pro74.4
GLM-4.774.3
DeepSeek-V4-Flash73.2
DeepSeek-V3.273.0
MiniMax M372.9
Opus 4.872.9
Sonnet 4.672.8
MiMo-V2.572.6
Mistral Small 472.4
DeepSeek-R1-052872.1
Trinity Large Preview70.8
GPT-5.5 Instant69.9
Grok 4.2068.3
Sonnet 567.5
Trinity Large Thinking66.8
GPT-4o66.6
Loading Atlas data…