Atlas

Benchmarks

← All benchmarks

MuSR (Open LLM Leaderboard v2)

General QA · 2024-06-26

MuSR poses multi-step soft reasoning problems over long narratives. The Open LLM Leaderboard v2 figure averages three subtasks with different answer-option counts, so it is a mean of means with no single guessing floor. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.

Top models (higher is better)

ModelScore
DeepSeek R1 Distill Qwen 14B53.7
DeepSeek LLM 67B Chat50.6
Phi-450.3
Neural Chat 7B v3-149.8
Qwen2.5 32B49.8
Hermes 3 Llama 3.1 70B49.5
Llama 3.1 Tülu 3 70B DPO49.2
C4AI Command R+ 08-202448.3
C4AI Command R+48.1
Qwen2.5 72B47.7
Stable Beluga 247.3
Qwen2-72B47.0
Llama 3.1 70B Instruct45.8
Llama 3.1 70B45.7
Phi-3.5-MoE-instruct45.6
Qwen2 72B Instruct45.6
Phi-3-Small-8K-Instruct45.6
Mistral Large 2.1 (Instruct 2411)45.4
DeepSeek-R1-Distill-Qwen-32B45.3
Qwen1.5 110B Chat45.2
Command R45.2
Qwen2 VL 72B Instruct44.9
Smaug-72B-v0.144.7
Llama 3.3 70B Instruct44.6
Gemma 2 9B44.6
Solar Pro Instruct Preview44.2
OpenChat 3.5 121044.1
Qwen1.5 14B Chat44.0
Gemma 2 27B44.0
Yi 1.5 6B Chat43.9
Qwen2.5 Coder 32B Instruct43.9
Qwen2 VL 7B Instruct43.8
DeepSeek-R1-Distill-Llama-70B43.4
Aria43.4
Llama 3.1 Nemotron 70B Instruct HF43.3
Mixtral 8x7B43.2
Mixtral 8x22B Instruct43.1
Aya-23-35B43.1
Qwen2 1.5B Instruct42.9
Yi 1.5 34B Chat42.8
Loading Atlas data…