Atlas

Benchmarks

← All benchmarks

BRIDGE Medical (few-shot)

Science

BRIDGE is a clinical benchmark suite from Yale's YLab that runs one model over 87 real-world clinical tasks spanning coding, hospitalization and mortality prediction, EHR question answering, clinical note understanding, and multilingual clinical text. The headline figure averages per-task scores whose own metrics and option counts differ, so neither a single denominator nor one guessing floor is well defined. This row is the few-shot prompting view, a component of the zero-shot factor group: it re-measures the same tasks on the same models under in-context examples, and scores run several points higher than zero-shot as a result.

Top models (higher is better)

ModelScore
Gemini 1.5 Pro 00255.5
Gemma 4 31B IT54.9
Gemini 2.5 Flash53.4
Gemini 2.0 Flash 00153.3
Gemma 4 26B A4B IT53.2
DeepSeek-R151.4
Qwen3-Next-80B-A3B-Thinking50.9
Mistral Large 2.1 (Instruct 2411)50.7
Qwen3 Next 80B A3B Instruct50.6
Llama 3.1 70B Instruct50.5
Qwen2.5 Instruct 32B49.5
Llama 3.3 70B Instruct49.5
Gemma 3 27B IT49.5
Magistral Small 1.049.1
Qwen3 32B48.9
Qwen3-235B-A22B48.7
QwQ-32B-Preview48.4
QwQ-32B48.2
Gemma 3 12B IT47.7
Qwen3-30B-A3B47.4
Phi-446.8
Qwen3 14B46.8
DeepSeek-R1-Distill-Llama-70B46.2
Gemma 2 27B IT45.8
Qwen3 8B45.5
DeepSeek-R1-Distill-Qwen-32B44.3
GPT-3.5 Turbo 012543.6
Llama 3.1 8B Instruct43.5
Gemma 2 9B IT43.5
DeepSeek R1 0528 Qwen3 8B42.9
Llama 3.1 Nemotron 70B Instruct HF42.8
Qwen3 4B42.8
K2-Think42.7
Qwen2.5 7B Instruct41.6
DeepSeek R1 Distill Qwen 14B41.4
Mistral Small 3.1 24B Instruct 250340.9
Llama 4 Scout Instruct40.6
Mistral Small 3 24B Instruct 250139.7
Ministral 8B Instruct 241039.6
gpt-oss-120b39.0
Loading Atlas data…