SM Bench — EQ Boundaries
Safety · 2026-02-01
SM Bench EQ Boundaries tests emotionally intelligent responses that remain supportive and proportionate without unnecessary clinical distancing, canned disclaimers, or inappropriate dependency reinforcement. Scores are difficulty-weighted judge credit across 100 fixed prompts.
Top models (higher is better)
| Model | Score |
|---|---|
| Grok 4.20 | 90.5 |
| Mistral Small 4 | 85.7 |
| MiMo-V2-Pro | 79.8 |
| Grok 4.1 Fast | 74.7 |
| MiMo-V2-Omni | 74.2 |
| GLM-5.1 | 73.9 |
| GPT-4.1 | 73.3 |
| GLM-4.7 | 71.3 |
| GPT-4o Mini | 71.3 |
| Grok 4.3 | 70.8 |
| Trinity Large Preview | 70.8 |
| MiMo-V2.5 | 70.5 |
| DeepSeek-V4-Flash | 69.9 |
| Grok 4.5 | 69.1 |
| GPT-5.3-Codex | 69.1 |
| GLM 5 Turbo | 69.1 |
| MiniMax M2.1 | 68.8 |
| Gemini 3.5 Flash | 68.5 |
| Gemini 3.1 Pro Preview | 68.3 |
| GPT-5.4 Mini | 68.3 |
| MiMo-V2.5-Pro | 68.0 |
| Gemma 4 31B IT | 67.7 |
| DeepSeek-V3.2 | 67.4 |
| Gemini 3.1 Flash-Lite Preview | 66.8 |
| GPT-5.6 Terra Pro | 66.6 |
| Gemini 3 Flash Preview | 65.5 |
| GPT-5.6 Terra | 65.5 |
| GPT-5.6 Sol Pro | 65.5 |
| DeepSeek-V4-Pro | 65.2 |
| GPT-5.6 Sol | 65.2 |
| GPT-5.3 Instant | 64.9 |
| Gemini 3 Pro Preview | 64.6 |
| GPT-5.5 | 64.6 |
| GLM-5.2 | 64.0 |
| Qwen3.7-Max | 63.5 |
| GLM-5 | 63.2 |
| MiMo-V2-Flash | 63.2 |
| Trinity Large Thinking | 62.9 |
| GPT-5.5 Instant | 62.4 |
| Qwen3.5 Plus (2026-02-15) | 62.1 |