CCP-Narrative-Bench - English
Safety · 2025-09-30
The English-prompted condition of CCP-Narrative-Bench measures the average percentage of rubric-listed CCP narratives reflected in model responses. Lower scores indicate that responses echo fewer of the benchmark's identified narratives.
Top models (lower is better)
| Model | Score |
|---|---|
| DeepSeek-R1 | 0.9 |
| GPT-5 | 1.4 |
| gpt-oss-120b | 1.4 |
| Opus 4 | 3.0 |
| DeepSeek-V3.1 | 5.3 |
| DeepSeek-R1-0528 | 15.9 |