Atlas

Benchmarks

← All benchmarks

HealthBench Professional

Professional Work · 2026-04-30

HealthBench Professional evaluates real clinical work using 525 physician-authored conversations spanning care consultations, writing and documentation, and medical research, each graded against rubric criteria by an LLM judge. This is the raw rubric score, before the benchmark's published length adjustment.

Top models (higher is better)

ModelScore
Claude Opus 573.4
Claude Mythos 570.3
GPT-5.6 Sol64.1
Sonnet 562.4
GPT-5.6 Terra62.4
Opus 4.860.3
GPT-5.6 Luna59.8
GPT-5.557.2
GPT-5.451.9
GPT-551.0
GPT-5.250.0
GPT-5.148.0
GPT-5.5 Instant40.7
GPT-5.1 Instant40.4
GPT-5.2 Instant38.3
GPT-5.3 Instant33.8
Loading Atlas data…