Atlas

Benchmarks

← All benchmarks

SM Bench — NSFW (System Prompt)

Safety · 2026-02-01

SM Bench NSFW (System Prompt) repeats the mature-content capability test under an explicit system-level instruction, exposing differences between default behaviour and system-instructed behaviour. Scores are difficulty-weighted judge credit across 100 fixed prompts.

Top models (higher is better)

ModelScore
Grok 4.5100.0
Grok 4.3100.0
Gemini 3 Flash Preview100.0
Grok 4.1 Fast100.0
GLM-5.1100.0
Opus 4.7100.0
Kimi K2.5100.0
Gemma 4 31B IT100.0
MiMo-V2-Pro100.0
Opus 4.6100.0
Gemini 3 Pro Preview100.0
MiMo-V2-Omni100.0
MiMo-V2.5-Pro100.0
DeepSeek-V4-Flash100.0
GLM-5.299.8
GPT-4.199.3
Gemini 3.1 Flash-Lite Preview99.1
GLM-599.1
Trinity Large Thinking99.1
Gemini 3.5 Flash98.6
Qwen3.7-Max98.6
Qwen3.6 Plus Preview98.6
DeepSeek-V3.298.6
DeepSeek-R1-052898.6
Claude Fable 598.6
DeepSeek-V4-Pro97.7
Sonnet 4.597.7
Kimi K2.696.8
Trinity Large Preview96.8
MiMo-V2.595.9
Gemini 3.1 Pro Preview95.4
Nemotron 3 Super 120B A12B93.3
Opus 4.592.4
MiniMax M392.2
Mistral Small 491.7
MiniMax M2.791.1
Sonnet 4.690.8
GPT-5.3 Instant89.0
Nemotron 3 Nano 30B A3B88.3
GLM-4.785.8
Loading Atlas data…