Atlas

Benchmarks

← All benchmarks

SM Bench — Creative Writing (Mature Themes)

Safety · 2026-02-01

SM Bench Creative Writing (Mature Themes) measures whether models can follow lawful creative-writing requests involving mature or sensitive framing without policy-shaped evasion. Scores are difficulty-weighted judge credit across 100 fixed prompts.

Top models (higher is better)

ModelScore
Grok 4.5100.0
Grok 4.3100.0
Gemini 3 Flash Preview100.0
Grok 4.1 Fast100.0
Gemini 3.1 Flash-Lite Preview100.0
DeepSeek-V4-Pro100.0
GPT-4.1100.0
DeepSeek-V3.2100.0
Mistral Small 499.1
Gemma 4 31B IT98.7
Gemini 3.5 Flash98.7
Qwen3.7-Max98.7
GLM-4.798.7
GLM-5.298.2
GLM-5.197.8
GPT-4o Mini97.8
Gemini 3.1 Pro Preview97.3
Kimi K2.597.3
MiMo-V2-Pro97.3
Gemini 3 Pro Preview97.3
GLM-597.3
GPT-4o97.3
Qwen3.6 Plus Preview96.0
DeepSeek-R1-052896.0
GPT-5.5 Instant96.0
GPT-5.6 Luna96.0
MiMo-V2-Omni94.6
GPT-5.6 Sol94.6
GPT-5.6 Terra93.7
DeepSeek-V4-Flash93.3
Grok 4.2093.3
Opus 4.792.8
Nemotron 3 Nano 30B A3B92.4
Trinity Large Preview91.3
Claude Fable 590.6
Opus 4.589.7
Kimi K2.689.7
Trinity Large Thinking89.7
Sonnet 589.2
GPT-5.6 Terra Pro88.3
Loading Atlas data…