HealthBench Hard
Professional Work · 2025-05-12
The 1,000-conversation hard subset of OpenAI's HealthBench, selected so that no frontier model at release scored well on it, and graded against the same conversation-specific physician-written rubrics. Ships in openai/simple-evals as healthbench_hard. This row holds scores from OpenAI's original HealthBench implementation, as reported up to March 2026.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.2 | 42.0 |
| GPT-5.4 | 40.1 |
| GPT-5.2 Instant | 26.8 |
| GPT-5.3 Instant | 25.9 |