MMLU Pro
General QA · 2024-06-03
Accuracy on the 14-subject MMLU Pro academic benchmark.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3 Pro Preview | 89.8 |
| Opus 4.5 | 89.5 |
| Gemini 3 Flash Preview | 88.2 |
| Opus 4.1 | 88.0 |
| MiniMax M2.1 | 87.5 |
| GPT-5.2 | 87.4 |
| Opus 4 | 87.3 |
| Grok 4 | 86.6 |
| GPT-5-Codex | 86.5 |
| DeepSeek-V3.2-Speciale | 86.3 |
| DeepSeek-V3.2 | 86.2 |
| GPT-5.1-Codex | 86.0 |
| Sonnet 4.5 | 86.0 |
| Gemini 2.5 Pro | 85.8 |
| Doubao-Seed-Code | 85.4 |
| Grok 4.1 Fast | 85.4 |
| o3 | 85.3 |
| DeepSeek-V3.1-Terminus | 85.1 |
| DeepSeek-V3.2-Exp | 85.0 |
| Grok 4 Fast | 85.0 |
| Cogito v2.1 671B | 84.9 |
| Kimi K2 Thinking | 84.8 |
| DeepSeek-R1 | 84.4 |
| MiMo-V2-Flash | 84.3 |
| Gemini 2.5 Flash Preview (09-2025) | 84.2 |
| Qwen3-Max (2025-09-23) | 84.1 |
| O1 | 84.1 |
| Qwen3 Max Preview | 83.8 |
| K-EXAONE 236B-A23B | 83.8 |
| Gemini 2.5 Pro Preview 05-06 | 83.7 |
| Claude 3.7 Sonnet | 83.7 |
| Qwen3 VL 235B A22B Thinking | 83.6 |
| GLM-4.5 | 83.5 |
| O4 Mini | 83.2 |
| Gemini 2.5 Flash | 83.2 |
| ERNIE 5.0 | 83.0 |
| Nova 2 Pro (Preview) | 83.0 |
| GLM 4.6 | 82.9 |
| Hermes 4 405B | 82.9 |
| Qwen3-235B-A22B | 82.8 |