Atlas

Benchmarks

← All benchmarks

MMLU Pro (Vals 5-Shot CoT)

General QA · 2024-06-03

Vals AI evaluates the full 14-subject MMLU-Pro benchmark using five in-context examples per subject and a chain-of-thought prompt. Its reported overall score is the unweighted mean of the 14 subject accuracies, rather than a micro-average over all 12,032 questions, so it is represented as an alternate component of the canonical MMLU-Pro factor family instead of independent evidence.

Top models (higher is better)

ModelScore
Claude Opus 591.6
Claude Fable 591.5
Gemini 3.1 Pro Preview91.0
Gemini 3 Pro Preview90.1
Opus 4.789.9
Opus 4.889.6
Gemini 3.5 Flash89.5
Qwen3.7-Max89.3
Gemini 3.6 Flash89.3
Grok 4.589.2
Opus 4.689.1
GPT-5.6 Sol89.1
Muse Spark 1.188.7
Gemini 3 Flash Preview88.6
GPT-5.588.1
Kimi K388.0
Opus 4.187.9
Qwen3.6 Plus (2026-04-02)87.7
Kimi K2.687.6
Sonnet 587.5
GPT-5.487.5
Sonnet 4.587.4
Sonnet 4.687.3
Muse Spark87.3
Opus 4.587.3
DeepSeek-V4-Pro87.2
Qwen3.5 Plus (2026-02-15)87.2
MiniMax M2.187.0
GLM-5.186.9
GLM-5.286.7
GPT-5.6 Terra86.7
GPT-586.5
GPT-5.186.4
Inkling86.3
Grok 4.2086.3
Gemini 3.1 Flash-Lite Preview86.2
GPT-5.286.2
Opus 486.2
GPT-5.6 Luna86.0
GLM-586.0
Loading Atlas data…