Atlas

Benchmarks

← All benchmarks

HealthBench Hard

Professional Work · 2025-05-12

The 1,000-conversation hard subset of OpenAI's HealthBench, selected so that no frontier model at release scored well on it, and graded against the same conversation-specific physician-written rubrics. Ships in openai/simple-evals as healthbench_hard. This row holds scores from OpenAI's original HealthBench implementation, as reported up to March 2026.

Top models (higher is better)

ModelScore
GPT-5.242.0
GPT-5.440.1
GPT-5.2 Instant26.8
GPT-5.3 Instant25.9
Loading Atlas data…