Atlas

Benchmarks

← All benchmarks

MMLU Medical Genetics (Open Medical-LLM)

Science · 2024-04-19

The medical-genetics subset of MMLU, 100 questions, scored by the Open Medical-LLM Leaderboard under lm-evaluation-harness with a fixed prompt. At 100 items the binomial noise floor alone is about 4 percentage points, so small differences here are not meaningful.

Top models (higher is better)

ModelScore
GPT-491.0
Llama 3 8B83.0
Llama 3 8B Instruct82.0
Solar 10.7B Instruct V1.076.0
GPT-3.5 Turbo 110674.0
Mistral 7B v0.1 Base71.0
Gemma 7B70.0
Qwen1.5 7B69.0
Qwen1.5 7B Chat67.0
Zephyr 7B Beta66.0
Gemma 1.1 7B IT66.0
Mistral 7B Instruct v0.163.0
Gemma 7B IT60.0
Phi-1.542.0
Falcon 7B30.0
Gemma 2B28.0
Loading Atlas data…