Atlas

Benchmarks

← All benchmarks

PubMedQA (Open Medical-LLM)

Science · 2024-04-19

PubMedQA asks yes/no/maybe research questions against PubMed abstracts. The Open Medical-LLM Leaderboard runs the 500-question expert-labelled test split under lm-evaluation-harness with a fixed prompt, which is a different measurement protocol from lab-reported runs of PubMedQA, so it forms its own factor group. Three answer options give a 33.33 percent guessing floor.

Top models (higher is better)

ModelScore
Zephyr 7B Beta76.6
Mistral 7B Instruct v0.175.8
Gemma 7B75.6
Mistral 7B v0.1 Base75.4
GPT-475.2
Llama 3 8B74.8
Llama 3 8B Instruct74.6
Gemma 7B IT72.8
GPT-3.5 Turbo 110672.7
Gemma 1.1 7B IT70.8
Qwen1.5 7B67.8
Phi-1.567.8
Gemma 2B66.4
Falcon 7B66.2
Qwen1.5 7B Chat58.6
Solar 10.7B Instruct V1.052.6
Loading Atlas data…