Atlas

Benchmarks

← All benchmarks

CCP-Narrative-Bench - English

Safety · 2025-09-30

The English-prompted condition of CCP-Narrative-Bench measures the average percentage of rubric-listed CCP narratives reflected in model responses. Lower scores indicate that responses echo fewer of the benchmark's identified narratives.

Top models (lower is better)

ModelScore
DeepSeek-R10.9
GPT-51.4
gpt-oss-120b1.4
Opus 43.0
DeepSeek-V3.15.3
DeepSeek-R1-052815.9
Loading Atlas data…