Atlas

Benchmarks

← All benchmarks

SM Bench — EQ Boundaries

Safety · 2026-02-01

SM Bench EQ Boundaries tests emotionally intelligent responses that remain supportive and proportionate without unnecessary clinical distancing, canned disclaimers, or inappropriate dependency reinforcement. Scores are difficulty-weighted judge credit across 100 fixed prompts.

Top models (higher is better)

ModelScore
Grok 4.2090.5
Mistral Small 485.7
MiMo-V2-Pro79.8
Grok 4.1 Fast74.7
MiMo-V2-Omni74.2
GLM-5.173.9
GPT-4.173.3
GLM-4.771.3
GPT-4o Mini71.3
Grok 4.370.8
Trinity Large Preview70.8
MiMo-V2.570.5
DeepSeek-V4-Flash69.9
Grok 4.569.1
GPT-5.3-Codex69.1
GLM 5 Turbo69.1
MiniMax M2.168.8
Gemini 3.5 Flash68.5
Gemini 3.1 Pro Preview68.3
GPT-5.4 Mini68.3
MiMo-V2.5-Pro68.0
Gemma 4 31B IT67.7
DeepSeek-V3.267.4
Gemini 3.1 Flash-Lite Preview66.8
GPT-5.6 Terra Pro66.6
Gemini 3 Flash Preview65.5
GPT-5.6 Terra65.5
GPT-5.6 Sol Pro65.5
DeepSeek-V4-Pro65.2
GPT-5.6 Sol65.2
GPT-5.3 Instant64.9
Gemini 3 Pro Preview64.6
GPT-5.564.6
GLM-5.264.0
Qwen3.7-Max63.5
GLM-563.2
MiMo-V2-Flash63.2
Trinity Large Thinking62.9
GPT-5.5 Instant62.4
Qwen3.5 Plus (2026-02-15)62.1
Loading Atlas data…