HealthBench (length-adjusted)
Professional Work · 2026-04-23
OpenAI's healthcare benchmark of 5,000 multi-turn conversations, scored against 48,562 conversation-specific rubric criteria written by 262 physicians and applied by a model grader, with the published length adjustment applied. Responses of 2,000 characters receive no adjustment; longer responses are penalised 2.99 points per additional 500 characters on the 0-100 scale and shorter responses receive the corresponding credit, so that verbosity cannot inflate the score by satisfying more positive rubric criteria. Models are not told about the penalty. OpenAI reports this as the headline HealthBench figure from the GPT-5.5 system card onward, quoting the unadjusted score alongside it.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 57.7 |
| GPT-5.6 Sol | 57.0 |
| GPT-5.6 Terra | 57.0 |
| GPT-5.2 | 56.8 |
| GPT-5.5 | 56.5 |
| GPT-5.6 Luna | 55.8 |
| GPT-5.4 | 54.0 |
| GPT-5.5 Instant | 51.4 |
| GPT-5.1 | 50.9 |
| GPT-5.2 Instant | 50.6 |
| GPT-5.1 Instant | 49.6 |
| GPT-5.3 Instant | 49.6 |