HealthBench Professional
Professional Work · 2026-04-30
HealthBench Professional evaluates real clinical work using 525 physician-authored conversations spanning care consultations, writing and documentation, and medical research, each graded against rubric criteria by an LLM judge. This is the raw rubric score, before the benchmark's published length adjustment.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 73.4 |
| Claude Mythos 5 | 70.3 |
| GPT-5.6 Sol | 64.1 |
| Sonnet 5 | 62.4 |
| GPT-5.6 Terra | 62.4 |
| Opus 4.8 | 60.3 |
| GPT-5.6 Luna | 59.8 |
| GPT-5.5 | 57.2 |
| GPT-5.4 | 51.9 |
| GPT-5 | 51.0 |
| GPT-5.2 | 50.0 |
| GPT-5.1 | 48.0 |
| GPT-5.5 Instant | 40.7 |
| GPT-5.1 Instant | 40.4 |
| GPT-5.2 Instant | 38.3 |
| GPT-5.3 Instant | 33.8 |