PubMedQA (Open Medical-LLM)
Science · 2024-04-19
PubMedQA asks yes/no/maybe research questions against PubMed abstracts. The Open Medical-LLM Leaderboard runs the 500-question expert-labelled test split under lm-evaluation-harness with a fixed prompt, which is a different measurement protocol from lab-reported runs of PubMedQA, so it forms its own factor group. Three answer options give a 33.33 percent guessing floor.
Top models (higher is better)
| Model | Score |
|---|---|
| Zephyr 7B Beta | 76.6 |
| Mistral 7B Instruct v0.1 | 75.8 |
| Gemma 7B | 75.6 |
| Mistral 7B v0.1 Base | 75.4 |
| GPT-4 | 75.2 |
| Llama 3 8B | 74.8 |
| Llama 3 8B Instruct | 74.6 |
| Gemma 7B IT | 72.8 |
| GPT-3.5 Turbo 1106 | 72.7 |
| Gemma 1.1 7B IT | 70.8 |
| Qwen1.5 7B | 67.8 |
| Phi-1.5 | 67.8 |
| Gemma 2B | 66.4 |
| Falcon 7B | 66.2 |
| Qwen1.5 7B Chat | 58.6 |
| Solar 10.7B Instruct V1.0 | 52.6 |