CCP-Narrative-Bench - Chinese
Safety · 2025-09-30
The Chinese-prompted condition of CCP-Narrative-Bench measures the average percentage of rubric-listed CCP narratives reflected in model responses. Lower scores indicate that responses echo fewer of the benchmark's identified narratives.
Top models (lower is better)
| Model | Score |
|---|---|
| GPT-5 | 2.5 |
| Opus 4 | 3.5 |
| gpt-oss-120b | 3.6 |
| DeepSeek-R1 | 9.3 |
| DeepSeek-V3.1 | 11.4 |
| DeepSeek-R1-0528 | 25.7 |