Atlas

Benchmarks

← All benchmarks

BBH (Open LLM Leaderboard v2)

General QA · 2024-06-26

BIG-Bench Hard collects tasks that were difficult for contemporary models at release. The Open LLM Leaderboard v2 figure averages accuracy over a subset of BBH subtasks whose option counts differ, so neither a single denominator nor one guessing floor is well defined. Scored by HuggingFace's Open LLM Leaderboard v2 under lm-evaluation-harness with a fixed prompt template and no chain-of-thought, which is a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying benchmark, so it forms its own factor group rather than pooling with them. Raw accuracy is imported, not the leaderboard's baseline-normalized column.

Top models (higher is better)

ModelScore
Qwen2 72B Instruct69.8
Qwen2 VL 72B Instruct69.5
Llama 3.3 70B Instruct69.2
Llama 3.1 70B Instruct69.2
Qwen2.5 Instruct 32B69.1
Solar Pro Instruct Preview68.2
Qwen2.5 72B68.0
Qwen2.5 32B67.7
Hermes 3 Llama 3.1 70B67.6
Mistral Large 2.1 (Instruct 2411)67.5
QwQ-32B-Preview66.9
Phi-466.9
Qwen2.5 Coder 32B Instruct66.3
Qwen2-72B66.2
Gemma 2 27B IT64.5
Phi-3 Medium 4K Instruct64.1
Phi-3.5-MoE-instruct64.1
Qwen2.5 14B Instruct63.9
Phi-3 Medium 128K Instruct63.8
Llama 3.1 Nemotron 70B Instruct HF63.2
Llama 3.1 70B62.6
Mixtral 8x22B62.4
Phi-3-Small-8K-Instruct62.1
Qwen1.5 110B Chat61.8
Llama 3.1 Tülu 3 70B DPO61.5
Mixtral 8x22B Instruct61.2
Yi 1.5 34B Chat60.8
Qwen1.5 32B Chat60.7
C4AI Command R+ 08-202460.0
Smaug-72B-v0.160.0
Gemma 2 9B IT59.9
DeepSeek R1 Distill Qwen 14B59.1
Qwen2.5-Coder-14B58.6
Stable Beluga 258.2
C4AI Command R+58.2
Qwen1.5 32B57.2
Aria57.0
Phi-4-mini-instruct56.9
Phi-3-mini-4k-instruct56.8
Aya Expanse 32B56.5
Loading Atlas data…