Atlas

Benchmarks

← All benchmarks

FORTRESS Average Risk Score (ARS)

Safety · 2025-06-17

Average Risk Score (ARS) across adversarial national-security and public-safety prompts; each response's harm score is the percentage of instance-specific rubric questions answered yes by a majority of three judges, and ARS averages those harm scores across prompts.

Top models (lower is better)

ModelScore
gpt-oss-120b8.2
Opus 4.59.6
Muse Spark 1.112.4
Sonnet 4.512.8
Claude 3.5 Sonnet (June 2024)13.0
Opus 4.613.0
Opus 4.114.8
GPT-5.4 Pro14.8
GPT-5 Pro15.2
o316.0
GPT-5.516.3
GPT-5 Mini17.0
GPT-517.0
GPT-5.217.5
gpt-oss-20b17.6
Sonnet 418.1
Opus 4.818.2
O119.4
Muse Spark20.2
Llama 3.1 405B20.6
O4 Mini21.5
Opus 424.8
GPT-5.125.7
Gemini 3.1 Pro Preview29.8
o3-mini30.1
GPT-5.1 Instant30.4
Haiku 3.530.4
Claude 3.7 Sonnet38.0
Llama 4 Maverick Instruct40.1
Kimi K2.541.1
Gemini 3 Pro Preview41.7
Llama 3.1 70B44.2
Llama 3.3 70B Instruct44.8
GPT-4o47.2
GPT-4o Mini48.1
Gemini 1.5 Flash 00250.6
Gemini 3.1 Flash-Lite Preview51.1
GPT-4.153.0
Gemini 1.5 Pro53.9
Gemini 2.5 Pro Preview 03-2554.9
Loading Atlas data…