MMLU-Pro (Open LLM Leaderboard v2)
General QA · 2024-06-26
MMLU-Pro extends MMLU to ten answer options with expert review to reduce noise. Reported scores land on the 12,032-question lattice, confirming a single-run proportion over the full set. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.