Atlas

Benchmarks

← All benchmarks

HPCT Refusal (CAIS)

Safety

HPCT Refusal measures how often a model refuses the Human Pathogen Capabilities Test's 100 text-only questions about conceptual and practical work on high-concern human pathogens. SAFE uses the raw refusal rate as a bioweapons-assistance risk component, with higher refusal rates indicating safer behavior.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Sonnet 4.597.0
Opus 4.790.0
Opus 4.885.0
Sonnet 4.678.0
Muse Spark 1.174.0
Opus 4.670.0
Opus 4.552.0
GPT-5.435.0
O131.0
GPT-5.129.0
GPT-5.529.0
o329.0
GPT-5.4 Mini28.0
GPT-526.0
O4 Mini26.0
GPT-5.6 Terra25.0
GPT-5.6 Sol24.0
GPT-5.221.0
GPT-5.6 Luna18.0
Gemini 3.1 Pro Preview7.0
Gemini 3.5 Flash6.0
Gemini 3 Flash Preview6.0
GPT-5 Nano2.0
Haiku 4.50.0
DeepSeek-R10.0
DeepSeek-V3.20.0
DeepSeek-V4-Pro0.0
Gemini 2.5 Flash0.0
Gemini 2.5 Flash-Lite Preview 06-170.0
Gemini 2.5 Pro0.0
Gemini 3.1 Flash-Lite Preview0.0
GLM-5.10.0
GLM-5.20.0
GPT-4o (2024-11-20)0.0
GPT-5.4 Nano0.0
GPT-5 Mini0.0
Grok 40.0
Grok 4.1 Fast0.0
Grok 4.200.0
Loading Atlas data…