Atlas

Benchmarks

← All benchmarks

MEDIC (closed-ended)

Science

MEDIC is M42's clinical evaluation suite. This row is its closed-ended headline: the average over MMLU medical subsets, MMLU-Pro, MedMCQA, MedQA, USMLE and PubMedQA. Those constituents have different option counts and denominators, so the average admits neither a single lattice nor one guessing floor, and it is carried as a mean of means rather than a proportion. Scored under M42's own harness, a different measurement protocol from lab-reported and Artificial Analysis runs of the same underlying tests, so it forms its own factor group instead of pooling with them.

Top models (higher is better)

ModelScore
Llama 3.1 70B Instruct79.8
Llama 3.3 70B Instruct79.5
DeepSeek-V378.8
Llama 4 Maverick Instruct77.7
Llama 4 Scout Instruct77.0
DeepSeek-R1-Distill-Llama-70B75.7
Mistral Large 2 (Instruct 2407)75.0
Qwen2.5 72B75.0
Llama 3 70B Instruct74.0
Qwen2 72B Instruct72.6
Qwen3 32B71.0
DeepSeek-R1-Distill-Qwen-32B70.6
QwQ-32B-Preview69.8
Qwen3 14B69.0
Gemma 3 27B IT68.6
Qwen3-30B-A3B68.5
Phi-467.5
Llama 3.1 8B Instruct67.2
Qwen3-235B-A22B65.7
Llama 3.1 Nemotron 70B Instruct HF65.6
Llama 3.1 70B65.6
Aya Expanse 32B65.5
DeepSeek R1 Distill Qwen 14B65.0
Llama 3.2 3B Instruct63.2
Qwen2.5 7B Instruct60.0
Solar 10.7B Instruct V1.059.4
Qwen3 4B56.8
Qwen3 8B56.4
Mistral Large 3 675B Instruct 251253.2
Kimi K2 Thinking52.3
Falcon2-11B51.9
DeepSeek-V3.151.6
Yi 1.5 6B Chat50.5
Llama 3.1 8B50.1
Qwen2.5 3B Instruct49.6
Mistral 7B Instruct v0.349.5
DeepSeek R1 Distill Llama 8B48.4
gpt-oss-120b42.0
Llama 3.2 1B Instruct39.4
Phi-3.5 Mini Instruct39.0
Loading Atlas data…