Atlas

Benchmarks

← All benchmarks

MMLU Pro - Physics

General QA · 2024-06-03

The Physics split of MMLU Pro. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Fable 595.8
Claude Opus 594.8
Gemini 3.1 Pro Preview94.7
Opus 4.694.1
Gemini 3.5 Flash94.1
Muse Spark 1.193.8
Gemini 3 Pro Preview93.7
Opus 4.893.5
Qwen3.7-Max93.5
DeepSeek-V4-Pro93.5
Gemini 3.6 Flash93.5
Opus 4.793.4
Gemini 3 Flash Preview93.4
GPT-5.6 Sol92.8
Qwen3.5 Plus (2026-02-15)92.8
Qwen3.6 Plus (2026-04-02)92.8
Muse Spark92.7
Sonnet 4.692.6
Kimi K392.6
GLM-5.192.3
Grok 4.392.3
GLM-5.292.2
Nemotron 3 Ultra 550B A55B92.2
Kimi K2.692.1
Gemini 3.1 Flash-Lite Preview92.0
Grok 4.592.0
Inkling92.0
GPT-5.6 Luna91.8
Sonnet 4.591.7
GLM-591.7
Sonnet 591.5
GPT-5.591.5
Grok 4.2091.5
GPT-5.291.4
Qwen3.5-Flash91.4
Kimi K2.591.3
GPT-5.491.2
MiMo-V2.5-Pro91.2
Qwen3 Max Thinking (2026-01-23)91.2
GPT-5.6 Terra91.1
Loading Atlas data…