Atlas

Benchmarks

← All benchmarks

MEDIC (clinical summarization)

Science

MEDIC's clinical summarization track, scored by rubric on coverage, conformity, consistency and conciseness and reported as an overall percentage. The overall figure is a mean of judged per-dimension rubric scores, not a pass rate over countable items, so no denominator is declared. Observed scores span a narrow high band (roughly 55 to 88), which is characteristic of rubric-judged generation and means the family carries less discrimination than its range suggests.

Top models (higher is better)

ModelScore
DeepSeek-V387.8
GPT-4.1 Mini87.5
Qwen2 72B Instruct87.5
Phi-487.4
O4 Mini87.2
Mistral Large 2 (Instruct 2407)87.1
Qwen3 32B87.0
Qwen3-235B-A22B86.8
Qwen3-30B-A3B86.6
Qwen2.5 7B Instruct86.2
DeepSeek-R1-Distill-Qwen-32B86.2
Qwen3 14B86.1
DeepSeek R1 Distill Llama 8B85.8
GPT-4.185.8
DeepSeek R1 Distill Qwen 14B85.8
DeepSeek-R1-Distill-Llama-70B85.6
Qwen2.5 3B Instruct85.6
Mistral 7B Instruct v0.385.5
Aya Expanse 32B85.2
Qwen3 4B85.0
QwQ-32B-Preview84.8
Llama 3.1 Nemotron 70B Instruct HF84.7
Llama 4 Scout Instruct84.3
Llama 3.1 70B Instruct84.3
Qwen2.5 72B84.1
Llama 3.1 405B Instruct84.0
Kimi K2 Thinking83.8
Llama 3.1 8B Instruct83.8
Qwen3 1.7B83.7
Qwen3 8B83.7
Llama 3.2 3B Instruct83.5
Llama 3.3 70B Instruct83.2
Llama 3 70B Instruct83.1
DeepSeek-V3.183.0
Llama 4 Maverick Instruct82.9
Qwen3 0.6B82.4
DeepSeek-R1-Distill-Qwen-1.5B82.3
gpt-oss-20b82.1
Mistral Large 3 675B Instruct 251281.7
Phi-3.5 Mini Instruct81.6
Loading Atlas data…