HealthBench Hard (length-adjusted)
Professional Work · 2026-04-23
HealthBench Hard with the published length adjustment applied: responses of 2,000 characters receive no adjustment, longer responses are penalised 3.92 points per additional 500 characters on the 0-100 scale, and shorter responses receive the corresponding credit. OpenAI reports this alongside the unadjusted score from the GPT-5.5 system card onward. Recorded as a diagnostic so that models are not entered twice under two readings of one run.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 34.7 |
| GPT-5.2 | 34.3 |
| GPT-5.6 Sol | 33.1 |
| GPT-5.6 Terra | 32.7 |
| GPT-5.6 Luna | 32.0 |
| GPT-5.5 | 31.5 |
| GPT-5.4 | 29.1 |
| GPT-5.1 | 25.4 |
| GPT-5.2 Instant | 23.3 |
| GPT-5.5 Instant | 22.9 |
| GPT-5.1 Instant | 21.6 |
| GPT-5.3 Instant | 20.2 |