Atlas

Benchmarks

← All benchmarks

Spiral-Bench v1.0 Safety Score

Safety · 2025-08-15

Earlier safety benchmark for sycophancy and delusion reinforcement.

Top models (higher is better)

ModelScore
GPT-5 (2025-08-07)87.0
o386.1
gpt-oss-120b81.4
Sonnet 4.576.4
O4 Mini73.3
Kimi K2 Instruct73.0
GPT-5 Chat (2025-08-07)58.1
Gemini 2.5 Flash49.1
Llama 4 Maverick Instruct48.1
Claude 3.5 Sonnet (Oct 2024)47.7
Gemini 2.5 Pro43.5
ChatGPT-4o Latest (source-unspecified snapshot)42.3
Sonnet 441.1
DeepSeek-R1-052822.4
Loading Atlas data…