Atlas

Benchmarks

← All benchmarks

EQ-Bench v2

Chat & Writing · 2023-12-11

Emotional-intelligence benchmark based on dialogue emotion prediction.

Top models (higher is better)

ModelScore
Claude 3.5 Sonnet (June 2024)86.4
GPT-4 Turbo86.3
GPT-4 Turbo (1106 Preview)86.0
GPT-485.7
Mistral Large 1.085.2
Mistral Large 2 (Instruct 2407)85.0
GPT-4 061384.8
GPT-4 0125 Preview83.9
Qwen1.5 110B Chat83.7
Gemini 1.5 Pro 00283.5
GPT-4o83.5
DeepSeek-V2-Chat-062883.2
Llama 3.1 405B Instruct83.0
Qwen1.5 72B Chat82.8
Hermes 3 Llama 3.1 405B82.8
ChatGPT-4o Latest (September 2024 benchmark entry)82.5
Opus 382.2
Llama 3 70B Instruct82.1
Llama 3.2 90B Vision Instruct82.0
DeepSeek-V2.582.0
Qwen2 72B Instruct81.3
Mistral Small Instruct 240980.9
Qwen 72B Chat80.7
Smaug-Llama-3-70B-Instruct80.7
Gemma 2 27B IT80.5
o1 Preview80.5
Gemma 2 9B IT80.5
Claude 3 Sonnet80.5
Mistral Small 1.080.4
Qwen2.5 Instruct 32B79.9
Smaug-72B-v0.179.8
Qwen2.5 14B Instruct79.2
Qwen2.5 72B Instruct79.0
Mixtral 8x22B Instruct78.8
Solar Pro Instruct Preview78.5
WizardLM-2 8x22B77.9
Quyen Pro Max V0.177.2
Mistral NeMo Instruct 240777.1
Phi-3.5-MoE-instruct77.0
GPT-4o Mini76.9
Loading Atlas data…