Atlas

Benchmarks

← All benchmarks

HealthBench Hard (updated implementation)

Professional Work · 2026-04-23

HealthBench Hard re-scored under the updated HealthBench implementation OpenAI introduced with the GPT-5.5 system card in April 2026, under which it recomputed prior models. Not comparable with Hard scores OpenAI published before 2026-04-23.

Top models (higher is better)

ModelScore
GPT-541.6
GPT-5.141.4
GPT-5.238.9
GPT-5.6 Terra34.3
GPT-5.533.8
GPT-5.6 Luna31.4
GPT-5.6 Sol31.1
GPT-5.430.3
GPT-5.2 Instant23.5
GPT-5.1 Instant23.0
GPT-5.5 Instant21.3
GPT-5.3 Instant17.8
Loading Atlas data…