Atlas

Benchmarks

← All benchmarks

Global MMLU Lite

General QA · 2024-12-04

Global MMLU Lite is a Cohere Labs multilingual and multicultural knowledge benchmark. It reports overall, culturally sensitive, culturally agnostic, and language-specific accuracy.

Top models (higher is better)

ModelScore
Claude Fable 593.3
Gemini 3.1 Pro Preview93.2
Gemini 3 Pro Preview92.2
Opus 4.692.2
GPT-5.6 Sol91.8
Gemini 3 Flash Preview91.4
Opus 4.591.3
GPT-590.7
GPT-5.190.6
Sonnet 4.690.5
Gemini 2.5 Pro90.3
Qwen3.5 397B A17B90.0
Gemini 2.5 Pro Experimental 03-2589.8
GPT-5.289.8
GPT-5.1-Codex89.7
Grok 489.5
Grok 4.2089.5
GPT-5.4 Mini89.4
Gemini 3.5 Flash-Lite89.4
DeepSeek-V4-Pro89.3
Sonnet 4.589.3
GLM-5.289.2
Qwen3.5 122B A10B89.1
GPT-5.6 Luna88.7
Inkling88.7
Gemini 2.5 Pro Preview 05-0688.6
Gemini 2.5 Flash88.4
Gemini 2.5 Flash Preview 04-1788.4
DeepSeek-V4-Flash88.4
Kimi K2.688.4
DeepSeek-V3.2-Speciale88.4
GPT-5 Mini87.4
Inkling-Small Preview86.8
Qwen3.5-27B86.8
DeepSeek-V3.2-Exp86.7
Inkling-Small86.7
DeepSeek-V3.286.5
Qwen3.5 35B A3B86.4
DeepSeek-R1-052886.0
Grok 4 Fast85.9
Loading Atlas data…