Atlas

Benchmarks

← All benchmarks

MAGI-Hard

General QA · 2024-04-01

Hard subset of MMLU and AGIEval for top-end discrimination.

Top models (higher is better)

ModelScore
Llama 3.1 405B Instruct83.8
Gemini 1.5 Pro 00281.8
GPT-4o80.9
ChatGPT-4o Latest (September 2024 benchmark entry)80.6
Claude 3.5 Sonnet (June 2024)78.8
GPT-4 061377.8
Llama 3.2 90B Vision Instruct77.8
Qwen2.5 72B Instruct77.8
GPT-4 Turbo77.7
GPT-4 0125 Preview76.8
Qwen2.5 Instruct 32B76.7
Opus 376.5
Hermes 3 Llama 3.1 405B76.2
Qwen2 72B Instruct75.7
GPT-475.7
GPT-4 Turbo (1106 Preview)75.0
Qwen2.5 14B Instruct72.8
Mistral Large 2 (Instruct 2407)72.4
Solar Pro Instruct Preview70.8
Llama 3 70B Instruct68.0
Mistral Large 1.067.7
GPT-4o Mini67.5
Phi-3.5-MoE-instruct67.3
Smaug-Llama-3-70B-Instruct67.3
Phi-3 Medium 4K Instruct66.4
Qwen1.5 110B Chat66.1
Yi 1.5 34B Chat64.8
Phi-3-Small-8K-Instruct64.2
Gemma 2 27B IT64.1
Llama 3 8B Instruct63.8
Qwen1.5 72B Chat63.5
Nous Hermes 2 Yi 34B63.0
Mixtral 8x22B Instruct62.4
DeepSeek-V2.562.0
Claude 3 Sonnet61.0
Qwen1.5 32B Chat60.7
DeepSeek-V2-Chat-062860.6
Qwen 72B Chat60.4
Smaug-72B-v0.160.2
Quyen Pro Max V0.159.3
Loading Atlas data…