MuSR (Open LLM Leaderboard v2)
General QA · 2024-06-26
MuSR poses multi-step soft reasoning problems over long narratives. The Open LLM Leaderboard v2 figure averages three subtasks with different answer-option counts, so it is a mean of means with no single guessing floor. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.