BRIDGE Medical (chain-of-thought)
Science
BRIDGE is a clinical benchmark suite from Yale's YLab that runs one model over 87 real-world clinical tasks spanning coding, hospitalization and mortality prediction, EHR question answering, clinical note understanding, and multilingual clinical text. The headline figure averages per-task scores whose own metrics and option counts differ, so neither a single denominator nor one guessing floor is well defined. This row is the chain-of-thought prompting view, a component of the zero-shot factor group: it re-measures the same tasks on the same models while asking for explicit reasoning.