HealthBench (updated implementation)
Professional Work · 2026-04-23
OpenAI's HealthBench re-scored under the updated implementation the company introduced with the GPT-5.5 system card in April 2026, which it describes as "an updated implementation of HealthBench" and under which it recomputed the scores of prior models. Same 5,000 conversations and rubric design as the original, but the regraded values differ enough that the two cannot be compared: GPT-5.2 scores 63.3 under the original grader and 60.7 under this one at essentially identical response length, and GPT-5.3 Instant scores 54.1 against 47.9. Scores published by OpenAI before 2026-04-23, and evaluations run by third parties that do not state which implementation they used, stay on the original HealthBench row.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.1 | 64.2 |
| GPT-5 | 63.1 |
| GPT-5.2 | 60.7 |
| GPT-5.6 Terra | 58.7 |
| GPT-5.5 | 58.4 |
| GPT-5.4 | 55.7 |
| GPT-5.6 Sol | 55.6 |
| GPT-5.6 Luna | 55.4 |
| GPT-5.2 Instant | 51.5 |
| GPT-5.5 Instant | 50.9 |
| GPT-5.1 Instant | 50.8 |
| GPT-5.3 Instant | 47.9 |