Atlas

Benchmarks

← All benchmarks

MMLU

General QA · 2020-09-07

MMLU (Measuring Massive Multitask Language Understanding) is a multiple-choice benchmark spanning 57 subjects across STEM, the humanities, the social sciences, and more. It measures broad world knowledge and problem-solving, typically reported as average accuracy. Higher scores indicate better performance.

Top models (higher is better)

ModelScore
Claude 3.5 Sonnet (June 2024)88.7
Claude 3.5 Sonnet (Oct 2024)88.7
Opus 388.2
GPT-4o (2024-11-20)88.1
Gemini 1.5 Pro 00286.9
GPT-486.4
Gemini 1.5 Pro 00185.9
Qwen2.5 72B85.0
Phi-484.8
Llama 3.1 405B Instruct84.5
Llama 3.1 405B84.4
GPT-4o (2024-08-06)84.3
GPT-4o84.2
GPT-4 061382.4
Qwen2 72B Instruct82.4
Nova Pro82.0
Claude 3 Sonnet81.5
Llama 3.2 90B Vision Instruct80.3
Llama 3.1 70B Instruct80.1
Qwen2.5 14B Instruct79.9
Gemini 2.0 Flash Experimental79.7
Llama 3 70B Instruct79.3
Yi-Large79.3
Claude 278.5
DeepSeek-V278.4
PaLM 2-L78.4
Jamba 1.5 Large78.2
Mixtral 8x22B77.8
Haiku 3.577.6
Claude 1.377.0
Nova Lite77.0
Qwen1.5 110B Chat76.8
Claude 3 Haiku76.7
Gemma 2 27B IT75.7
Phi-3-Small-8K-Instruct75.7
Qwen2.5-Coder-14B75.2
Qwen1.5 32B74.4
DBRX Instruct74.1
Gemini 1.5 Flash 00273.9
Claude 2.173.5
Loading Atlas data…