Atlas

Benchmarks

← All benchmarks

BioSecBench-Refusal

Safety · 2026-06-30

Biosecurity refusal and over-refusal benchmark testing whether agents block genuinely harmful dual-use biology requests while still helping with benign research. The headline balanced score is the trial-weighted harmonic mean of red-team refusal rate and routine compliance rate, so a model must both refuse hazardous tasks and allow legitimate ones.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash51.5
GPT-5.445.7
Sonnet 4.643.5
Gemini 3.1 Pro Preview43.0
GPT-5.6 Sol39.5
Opus 4.639.0
GPT-5.537.1
Opus 4.736.7
GPT-5.6 Terra34.3
GPT-5.6 Luna33.3
Sonnet 530.1
Opus 4.828.6
Grok 4.312.9
Grok 4.59.3
Grok 4.203.1
Loading Atlas data…