Global MMLU Lite
General QA · 2024-12-04
Global MMLU Lite is a Cohere Labs multilingual and multicultural knowledge benchmark. It reports overall, culturally sensitive, culturally agnostic, and language-specific accuracy.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 93.3 |
| Gemini 3.1 Pro Preview | 93.2 |
| Gemini 3 Pro Preview | 92.2 |
| Opus 4.6 | 92.2 |
| GPT-5.6 Sol | 91.8 |
| Gemini 3 Flash Preview | 91.4 |
| Opus 4.5 | 91.3 |
| GPT-5 | 90.7 |
| GPT-5.1 | 90.6 |
| Sonnet 4.6 | 90.5 |
| Gemini 2.5 Pro | 90.3 |
| Qwen3.5 397B A17B | 90.0 |
| Gemini 2.5 Pro Experimental 03-25 | 89.8 |
| GPT-5.2 | 89.8 |
| GPT-5.1-Codex | 89.7 |
| Grok 4 | 89.5 |
| Grok 4.20 | 89.5 |
| GPT-5.4 Mini | 89.4 |
| Gemini 3.5 Flash-Lite | 89.4 |
| DeepSeek-V4-Pro | 89.3 |
| Sonnet 4.5 | 89.3 |
| GLM-5.2 | 89.2 |
| Qwen3.5 122B A10B | 89.1 |
| GPT-5.6 Luna | 88.7 |
| Inkling | 88.7 |
| Gemini 2.5 Pro Preview 05-06 | 88.6 |
| Gemini 2.5 Flash | 88.4 |
| Gemini 2.5 Flash Preview 04-17 | 88.4 |
| DeepSeek-V4-Flash | 88.4 |
| Kimi K2.6 | 88.4 |
| DeepSeek-V3.2-Speciale | 88.4 |
| GPT-5 Mini | 87.4 |
| Inkling-Small Preview | 86.8 |
| Qwen3.5-27B | 86.8 |
| DeepSeek-V3.2-Exp | 86.7 |
| Inkling-Small | 86.7 |
| DeepSeek-V3.2 | 86.5 |
| Qwen3.5 35B A3B | 86.4 |
| DeepSeek-R1-0528 | 86.0 |
| Grok 4 Fast | 85.9 |