Atlas

Benchmarks

← All benchmarks

HealthBench Hard (length-adjusted)

Professional Work · 2026-04-23

HealthBench Hard with the published length adjustment applied: responses of 2,000 characters receive no adjustment, longer responses are penalised 3.92 points per additional 500 characters on the 0-100 scale, and shorter responses receive the corresponding credit. OpenAI reports this alongside the unadjusted score from the GPT-5.5 system card onward. Recorded as a diagnostic so that models are not entered twice under two readings of one run.

Top models (higher is better)

ModelScore
GPT-534.7
GPT-5.234.3
GPT-5.6 Sol33.1
GPT-5.6 Terra32.7
GPT-5.6 Luna32.0
GPT-5.531.5
GPT-5.429.1
GPT-5.125.4
GPT-5.2 Instant23.3
GPT-5.5 Instant22.9
GPT-5.1 Instant21.6
GPT-5.3 Instant20.2
Loading Atlas data…