Atlas

Benchmarks

← All benchmarks

BRIDGE Medical (zero-shot)

Science

BRIDGE is a clinical benchmark suite from Yale's YLab that runs one model over 87 real-world clinical tasks spanning coding, hospitalization and mortality prediction, EHR question answering, clinical note understanding, and multilingual clinical text, in the language each dataset was collected in. The headline figure averages per-task scores whose own metrics and option counts differ, so neither a single denominator nor one guessing floor is well defined. This row is the zero-shot prompting view, the suite's default protocol. The three prompting views form one factor group with zero-shot as the primary, because they re-measure the same tasks on the same models rather than measuring different abilities.

Top models (higher is better)

ModelScore
Gemma 4 31B IT46.7
Gemini 2.5 Flash44.8
Gemma 4 26B A4B IT44.5
DeepSeek-R144.3
Qwen3-Next-80B-A3B-Thinking43.9
Gemini 1.5 Pro 00243.9
Gemini 2.0 Flash 00143.0
Mistral Large 2.1 (Instruct 2411)42.3
Qwen3-235B-A22B41.6
Qwen3 32B41.0
Qwen3-30B-A3B40.9
Qwen3 14B40.2
Qwen3 8B40.0
Gemma 3 27B IT39.9
Qwen2.5 Instruct 32B39.9
Llama 3.3 70B Instruct39.9
Qwen3 Next 80B A3B Instruct39.8
DeepSeek-R1-Distill-Llama-70B39.8
DeepSeek-R1-Distill-Qwen-32B39.8
Mistral Small 3.1 24B Instruct 250339.7
QwQ-32B39.4
Llama 3.1 70B Instruct39.1
Qwen3 4B38.5
Gemma 2 27B IT38.2
K2-Think37.8
Magistral Small 1.037.6
Mistral Small 3 24B Instruct 250137.5
Gemma 3 12B IT37.3
gpt-oss-120b37.2
DeepSeek R1 0528 Qwen3 8B36.9
Phi-436.1
GPT-3.5 Turbo 012535.3
Mistral Small Instruct 240935.2
Llama 4 Scout Instruct35.1
Gemma 2 9B IT35.1
DeepSeek R1 Distill Qwen 14B34.3
Llama 3.1 Nemotron 70B Instruct HF32.8
QwQ-32B-Preview31.7
Qwen2.5 7B Instruct31.3
Ministral 8B Instruct 241030.4
Loading Atlas data…