Atlas

Benchmarks

← All benchmarks

MMMLU (CAISI 1,400-question subset)

General QA · 2025-09-30

CAISI's MMMLU subset contains 100 randomly selected questions from each of the benchmark's 14 translated languages, for 1,400 questions total. Questions span all source disciplines, and scores report multiple-choice accuracy.

Top models (higher is better)

ModelScore
GPT-587.7
Opus 483.8
DeepSeek-V3.182.2
DeepSeek-R1-052881.9
DeepSeek-R181.6
gpt-oss-120b77.7
Loading Atlas data…