Atlas

Benchmarks

← All benchmarks

BRIDGE Medical (chain-of-thought)

Science

BRIDGE is a clinical benchmark suite from Yale's YLab that runs one model over 87 real-world clinical tasks spanning coding, hospitalization and mortality prediction, EHR question answering, clinical note understanding, and multilingual clinical text. The headline figure averages per-task scores whose own metrics and option counts differ, so neither a single denominator nor one guessing floor is well defined. This row is the chain-of-thought prompting view, a component of the zero-shot factor group: it re-measures the same tasks on the same models while asking for explicit reasoning.

Top models (higher is better)

ModelScore
Gemma 4 31B IT45.9
Gemma 4 26B A4B IT45.2
Gemini 2.5 Flash43.3
Qwen3-Next-80B-A3B-Thinking42.9
DeepSeek-R142.1
Gemini 2.0 Flash 00142.0
Gemini 1.5 Pro 00240.5
Qwen3 Next 80B A3B Instruct40.5
Qwen3-235B-A22B40.1
Qwen3-30B-A3B39.4
DeepSeek-R1-Distill-Llama-70B39.0
Mistral Large 2.1 (Instruct 2411)38.9
DeepSeek-R1-Distill-Qwen-32B38.7
Qwen2.5 Instruct 32B37.7
Gemma 3 27B IT37.5
Qwen3 14B37.1
Qwen3 8B37.1
QwQ-32B37.0
Qwen3 4B37.0
Llama 3.3 70B Instruct36.8
DeepSeek R1 0528 Qwen3 8B36.3
Mistral Small 3.1 24B Instruct 250336.2
Gemma 3 12B IT35.4
K2-Think35.4
Llama 3.1 70B Instruct35.1
DeepSeek R1 Distill Qwen 14B34.8
Qwen3 32B34.5
Gemma 2 27B IT34.2
Magistral Small 1.034.2
Phi-432.6
gpt-oss-120b32.1
GPT-3.5 Turbo 012531.6
Mistral Small 3 24B Instruct 250131.6
Mistral Small Instruct 240931.2
Qwen2.5 7B Instruct30.3
Gemma 2 9B IT29.9
Llama 3.1 8B Instruct29.4
Llama 4 Scout Instruct29.4
Gemma 3 4B IT28.2
Qwen3 1.7B27.7
Loading Atlas data…