Atlas

Benchmarks

← All benchmarks

MEDIC (open-ended Elo)

Science

MEDIC's open-ended clinical track, scored by pairwise comparison and reported as an Elo rating with a 95% interval. Elo has no defined ceiling and its origin and unit are conventional, so no bounds are declared and the capability model reads it on its linear arm rather than through the bounded item curve. The companion 0-10 judge score is not imported: it is a monotone re-expression of the same comparisons, so carrying both would double-count one measurement.

Top models (higher is better)

ModelScore
gpt-oss-120b2032
Llama 3.1 Nemotron 70B Instruct HF1897
Qwen3-30B-A3B1874
gpt-oss-20b1833
Qwen3-235B-A22B1813
Qwen3 32B1813
GPT-5 Mini1790
O4 Mini1781
Qwen3 8B1773
Gemini 2.5 Flash Preview 04-171725
DeepSeek-V31678
Aya Expanse 32B1670
Llama 3.1 70B Instruct1666
Qwen3 14B1664
Qwen3 4B1653
Llama 3 70B Instruct1649
Llama 4 Scout Instruct1646
Llama 3.1 405B Instruct1643
Llama 4 Maverick Instruct1623
Llama 3.1 8B Instruct1607
GPT-4.11603
Llama 3.2 3B Instruct1587
Gemini 2.0 Flash1579
Phi-41557
GPT-4.1 Mini1542
DeepSeek-R1-Distill-Qwen-32B1503
DeepSeek R1 Distill Qwen 14B1501
DeepSeek-R1-Distill-Llama-70B1493
Mistral Large 2 (Instruct 2407)1482
Yi 1.5 6B Chat1461
Llama 3.2 1B Instruct1459
Qwen2 72B Instruct1438
DeepSeek R1 Distill Llama 8B1421
Qwen3 1.7B1417
Qwen2.5 7B Instruct1372
QwQ-32B-Preview1344
Mistral 7B Instruct v0.31328
Solar 10.7B Instruct V1.01273
Qwen2.5 72B1261
Qwen2.5 3B Instruct1249
Loading Atlas data…