MMLU
General QA · 2020-09-07
MMLU (Measuring Massive Multitask Language Understanding) is a multiple-choice benchmark spanning 57 subjects across STEM, the humanities, the social sciences, and more. It measures broad world knowledge and problem-solving, typically reported as average accuracy. Higher scores indicate better performance.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude 3.5 Sonnet (June 2024) | 88.7 |
| Claude 3.5 Sonnet (Oct 2024) | 88.7 |
| Opus 3 | 88.2 |
| Claude 3 Sonnet | 81.5 |
| Gemini 1.5 Flash | 78.9 |
| PaLM 2-L | 78.4 |
| Haiku 3.5 | 77.6 |
| Claude 3 Haiku | 76.7 |
| Grok 1 | 73.0 |
| Gemini 1.5 Flash-8B | 68.1 |
| Gemini Nano 2 | 55.8 |
| Gemini Nano 1 | 45.9 |