Atlas

Benchmarks

← All benchmarks

SM Bench — NSFW (Explicit)

Safety · 2026-02-01

SM Bench NSFW (Explicit) tests default model behaviour on explicit adult-content requests according to the benchmark's permissive usefulness rubric. Scores are difficulty-weighted judge credit across 100 fixed prompts.

Top models (higher is better)

ModelScore
Grok 4.5100.0
Grok 4.3100.0
Grok 4.1 Fast100.0
Kimi K2.580.3
DeepSeek-V4-Pro79.1
Gemini 3 Flash Preview66.7
DeepSeek-V3.264.2
GLM-5.162.2
Mistral Small 460.5
Kimi K2.655.2
Sonnet 4.554.2
Opus 4.754.0
Gemini 3.1 Pro Preview45.8
Opus 4.544.3
Opus 4.639.8
GLM-5.238.8
GLM 5 Turbo35.8
GLM-534.3
Grok 4.2030.4
DeepSeek-V4-Flash24.4
Gemini 3.1 Flash-Lite Preview24.1
GPT-4o Mini19.1
Trinity Large Preview18.9
Gemini 3 Pro Preview17.2
MiMo-V2-Omni17.2
Qwen3.7-Max16.9
Trinity Large Thinking15.4
GLM-4.714.4
Gemma 4 31B IT10.2
DeepSeek-R1-05289.9
Gemini 3.5 Flash8.7
MiniMax M38.0
Qwen3.6 Plus Preview7.0
MiniMax M2.57.0
Nemotron 3 Super 120B A12B7.0
Nemotron 3 Nano 30B A3B5.0
Sonnet 4.64.5
MiMo-V2-Flash4.5
GPT-4.14.2
MiniMax M2.74.0
Loading Atlas data…