Atlas

Benchmarks

← All benchmarks

SM Bench — Adversarial (Hostile Logic)

Safety · 2026-02-01

SM Bench Adversarial (Hostile Logic) tests resistance to manipulative framing, false premises, and hostile reasoning traps while preserving useful instruction following. Scores are difficulty-weighted judge credit across 100 fixed prompts.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash92.2
Gemini 3.1 Flash-Lite Preview91.7
Gemma 4 31B IT91.7
Gemini 3.1 Pro Preview90.7
GLM-590.7
GPT-5.5 Instant88.8
GLM-5.188.3
MiMo-V2-Pro88.3
GPT-4.187.3
GLM-5.286.8
Gemini 3 Pro Preview86.8
Qwen3.7-Max86.8
MiMo-V2-Omni86.6
GPT-5.186.3
Qwen3.6 Plus Preview85.8
GPT-5.6 Luna85.8
GPT-5.6 Luna Pro85.8
Grok 4.585.4
GPT-5.6 Sol Pro85.4
DeepSeek-R1-052884.9
GPT-5.6 Terra Pro84.9
Gemini 3 Flash Preview84.4
Kimi K2.584.4
Opus 4.583.9
Kimi K2.683.9
Grok 4.383.4
Opus 4.683.4
MiMo-V2.5-Pro83.4
GLM-4.782.9
GPT-5.582.9
GPT-5.6 Terra82.4
GPT-5 Mini82.4
DeepSeek-V4-Flash82.0
GPT-5.6 Sol82.0
Opus 4.781.5
Grok 4.1 Fast81.0
DeepSeek-V4-Pro81.0
Sonnet 581.0
Opus 4.880.5
MiMo-V2.580.5
Loading Atlas data…