MMMLU (CAISI 1,400-question subset)
General QA · 2025-09-30
CAISI's MMMLU subset contains 100 randomly selected questions from each of the benchmark's 14 translated languages, for 1,400 questions total. Questions span all source disciplines, and scores report multiple-choice accuracy.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 87.7 |
| Opus 4 | 83.8 |
| DeepSeek-V3.1 | 82.2 |
| DeepSeek-R1-0528 | 81.9 |
| DeepSeek-R1 | 81.6 |
| gpt-oss-120b | 77.7 |