Atlas

Benchmarks

← All benchmarks

MMLU Pro - Biology

General QA · 2024-06-03

The Biology split of MMLU Pro. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Opus 595.7
Gemini 3 Pro Preview95.4
Gemini 3.1 Pro Preview95.3
Qwen3.7-Max95.1
Muse Spark 1.195.0
Opus 4.894.7
Gemini 3.6 Flash94.7
Gemini 3 Flash Preview94.7
Opus 4.794.6
GPT-5.6 Sol94.6
GLM-594.3
Grok 4.594.3
Inkling94.3
Grok 4.2094.1
Qwen3.5 Plus (2026-02-15)94.1
Opus 4.594.0
Opus 4.694.0
GLM-5.294.0
GPT-5.594.0
Gemini 3.5 Flash93.9
Kimi K393.9
GLM-5.193.7
Kimi K2.693.7
Muse Spark93.7
Qwen3.6 Plus (2026-04-02)93.7
Claude Fable 593.6
Sonnet 593.6
MiniMax M2.193.6
GPT-5.6 Luna93.3
Grok 4.393.3
Gemini 3.1 Flash-Lite Preview93.2
Opus 493.0
GPT-5.6 Terra93.0
GPT-593.0
Grok 4.1 Fast93.0
Kimi K2.593.0
Qwen3.5-Flash93.0
Opus 4.192.9
DeepSeek-V4-Pro92.9
GPT-5.492.9
Loading Atlas data…