Atlas

Benchmarks

← All benchmarks

MMLU Pro - Engineering

General QA · 2024-06-03

The Engineering split of MMLU Pro. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Fable 590.0
Claude Opus 589.9
Muse Spark 1.188.1
Opus 4.788.0
Gemini 3.1 Pro Preview88.0
Gemini 3.5 Flash87.8
GPT-5.6 Sol87.8
Opus 4.887.7
Gemini 3 Pro Preview87.6
Gemini 3 Flash Preview87.3
Grok 4.587.2
Opus 4.687.1
Kimi K2.687.1
Kimi K386.9
DeepSeek-V4-Pro86.7
Gemini 3.6 Flash86.7
Opus 4.586.5
Qwen3.7-Max85.8
GPT-5.6 Luna85.7
Inkling85.7
Sonnet 4.685.2
GPT-5.485.2
Qwen3.5 Plus (2026-02-15)85.2
GPT-5.584.8
Muse Spark84.6
Qwen3.6 Plus (2026-04-02)84.3
Grok 4.384.2
GLM-5.184.0
Kimi K2.583.8
Gemini 3.1 Flash-Lite Preview83.7
GLM-5.283.4
DeepSeek-V3.283.3
GLM-583.2
Grok 4.2083.2
GPT-5.283.0
GPT-5.6 Terra82.9
Gemini 3.5 Flash-Lite82.8
Nemotron 3 Ultra 550B A55B82.8
Sonnet 4.582.5
Grok 482.5
Loading Atlas data…