Atlas
/
Benchmarks
137 sources · 855 models × 1278 benchmarks · 74,806 scores
☾
☀
Capabilities
Models
Compare
Benchmarks
Coverage
Methodology
← All benchmarks
MAGI-Hard
General QA · 2024-04-01
Hard subset of MMLU and AGIEval for top-end discrimination.
Top models
(higher is better)
Model
Score
Llama 3.1 405B Instruct
83.8
Gemini 1.5 Pro 002
81.8
GPT-4o
80.9
ChatGPT-4o Latest (September 2024 benchmark entry)
80.6
Claude 3.5 Sonnet (June 2024)
78.8
GPT-4 0613
77.8
Llama 3.2 90B Vision Instruct
77.8
Qwen2.5 72B Instruct
77.8
GPT-4 Turbo
77.7
GPT-4 0125 Preview
76.8
Qwen2.5 Instruct 32B
76.7
Opus 3
76.5
Hermes 3 Llama 3.1 405B
76.2
Qwen2 72B Instruct
75.7
GPT-4
75.7
GPT-4 Turbo (1106 Preview)
75.0
Qwen2.5 14B Instruct
72.8
Mistral Large 2 (Instruct 2407)
72.4
Solar Pro Instruct Preview
70.8
Llama 3 70B Instruct
68.0
Mistral Large 1.0
67.7
GPT-4o Mini
67.5
Phi-3.5-MoE-instruct
67.3
Smaug-Llama-3-70B-Instruct
67.3
Phi-3 Medium 4K Instruct
66.4
Qwen1.5 110B Chat
66.1
Yi 1.5 34B Chat
64.8
Phi-3-Small-8K-Instruct
64.2
Gemma 2 27B IT
64.1
Llama 3 8B Instruct
63.8
Qwen1.5 72B Chat
63.5
Nous Hermes 2 Yi 34B
63.0
Mixtral 8x22B Instruct
62.4
DeepSeek-V2.5
62.0
Claude 3 Sonnet
61.0
Qwen1.5 32B Chat
60.7
DeepSeek-V2-Chat-0628
60.6
Qwen 72B Chat
60.4
Smaug-72B-v0.1
60.2
Quyen Pro Max V0.1
59.3
Loading Atlas data…