Atlas

Benchmarks

← All benchmarks

MMLU Pro (CAISI science 1,000)

Science · 2025-09-30

CAISI's MMLU-Pro science subset contains 1,000 randomly selected questions from mathematics, physics, chemistry, engineering, biology, and computer science. All evaluated models receive the same selection, and scores report multiple-choice accuracy.

Top models (higher is better)

ModelScore
Opus 490.2
GPT-589.8
DeepSeek-R1-052889.7
DeepSeek-V3.189.0
DeepSeek-R187.5
gpt-oss-120b85.5
Loading Atlas data…