Atlas

Benchmarks

← All benchmarks

HealthBench Professional (Length-Adjusted)

Professional Work · 2026-04-30

HealthBench Professional evaluates real clinical work using 525 physician-authored conversations spanning care consultations, writing and documentation, and medical research. This metric applies the benchmark's published length adjustment to penalize unnecessarily verbose responses.

Top models (higher is better)

ModelScore
Claude Mythos 566.0
Claude Fable 560.9
GPT-5.6 Sol60.5
Claude Opus 559.8
Sonnet 557.8
GPT-5.6 Terra57.7
Opus 4.857.4
GPT-5.6 Luna55.7
GPT-5.448.1
GPT-546.2
GPT-5.245.9
GPT-5.139.6
GPT-5.5 Instant38.4
GPT-5.1 Instant37.6
GPT-5.2 Instant35.7
GPT-5.3 Instant32.9
Loading Atlas data…