Atlas

Benchmarks

← All benchmarks

EQ-Bench v2 + MAGI-Hard Combined

Chat & Writing · 2024-04-01

This legacy leaderboard metric averages EQ-Bench v2 and MAGI-Hard performance into one percentage score. Higher values indicate stronger combined performance.

Top models (higher is better)

ModelScore
Llama 3.1 405B Instruct83.4
Gemini 1.5 Pro 00282.7
Claude 3.5 Sonnet (June 2024)82.6
GPT-4o82.2
GPT-4 Turbo82.0
ChatGPT-4o Latest (September 2024 benchmark entry)81.5
GPT-4 061381.3
GPT-480.7
GPT-4 Turbo (1106 Preview)80.5
GPT-4 0125 Preview80.3
Llama 3.2 90B Vision Instruct79.9
Hermes 3 Llama 3.1 405B79.5
Opus 379.4
Mistral Large 2 (Instruct 2407)78.7
Qwen2 72B Instruct78.5
Qwen2.5 72B Instruct78.4
Qwen2.5 Instruct 32B78.3
Mistral Large 1.076.4
Qwen2.5 14B Instruct76.0
Llama 3 70B Instruct75.0
Qwen1.5 110B Chat74.9
Solar Pro Instruct Preview74.7
Smaug-Llama-3-70B-Instruct74.0
Qwen1.5 72B Chat73.1
Gemma 2 27B IT72.3
GPT-4o Mini72.2
Phi-3.5-MoE-instruct72.1
DeepSeek-V2.572.0
DeepSeek-V2-Chat-062871.9
Phi-3 Medium 4K Instruct71.4
Claude 3 Sonnet70.7
Mixtral 8x22B Instruct70.6
Qwen 72B Chat70.5
Smaug-72B-v0.170.0
Gemma 2 9B IT69.2
Yi 1.5 34B Chat68.9
Phi-3-Small-8K-Instruct68.8
WizardLM-2 8x22B68.5
Quyen Pro Max V0.168.2
Qwen1.5 32B Chat68.2
Loading Atlas data…