Atlas

Benchmarks

← All benchmarks

EQ-Bench 3: Compliant

Chat & Writing · 2025-04-28

Within EQ-Bench 3's emotional-intelligence scenarios, this subscore rates the compliant dimension of model responses on a 0–10 scale. Higher values are better.

Top models (higher is better)

ModelScore
Grok 3 Mini8.2
GPT-4.1 Mini8.1
Grok 38.1
Grok 47.7
Llama 3.1 Nemotron Ultra 253B V17.7
ChatGPT-4o (2025-04-25 update)7.6
Gemini 2.5 Flash Preview 05-207.6
GPT-4.1 Nano7.6
Grok 4.1 Fast7.6
GPT-4.17.5
O4 Mini7.5
Qwen2.5 72B Instruct7.5
GPT-47.4
GPT-4.57.4
GLM-4.57.3
ChatGPT-4o Latest observed 2025-03-277.2
Gemini 2.5 Pro Preview 05-067.2
Gemini 2.5 Pro Preview 03-257.2
Mistral Small 3 24B Instruct 25017.2
Qwen3 32B7.2
Llama 4 Maverick Instruct7.1
Qwen3-30B-A3B7.1
Llama 3.1 405B Instruct7.0
GPT-5 Chat (2025-08-07)6.9
Mistral Large 2.1 (Instruct 2411)6.9
Qwen3-235B-A22B6.9
Gemini 2.5 Pro Preview 06-056.8
Gemma 4 12B IT6.8
QwQ-32B6.8
gpt-oss-120b6.7
o36.7
Llama 4 Scout Instruct6.6
Mistral Small 3.1 24B Instruct 25036.6
Gemma 2 9B IT6.4
gpt-oss-20b6.3
Horizon Alpha6.3
Qwen3.5 397B A17B6.3
GPT-5.16.2
Claude 3.7 Sonnet6.1
DeepSeek-V4-Flash6.1
Loading Atlas data…