Atlas

Benchmarks

← All benchmarks

MMLU Professional Medicine (Open Medical-LLM)

Science · 2024-04-19

The professional-medicine subset of MMLU, 272 questions, scored by the Open Medical-LLM Leaderboard under lm-evaluation-harness with a fixed prompt.

Top models (higher is better)

ModelScore
GPT-493.8
Llama 3 8B Instruct74.6
Solar 10.7B Instruct V1.073.2
GPT-3.5 Turbo 110672.8
Llama 3 8B70.2
Mistral 7B v0.1 Base68.4
Qwen1.5 7B Chat65.4
Zephyr 7B Beta65.4
Gemma 7B63.2
Mistral 7B Instruct v0.159.6
Qwen1.5 7B59.6
Gemma 1.1 7B IT54.0
Gemma 7B IT47.1
Phi-1.529.0
Gemma 2B18.0
Falcon 7B17.6
Loading Atlas data…