Atlas

Benchmarks

← All benchmarks

CCP-Narrative-Bench - Chinese

Safety · 2025-09-30

The Chinese-prompted condition of CCP-Narrative-Bench measures the average percentage of rubric-listed CCP narratives reflected in model responses. Lower scores indicate that responses echo fewer of the benchmark's identified narratives.

Top models (lower is better)

ModelScore
GPT-52.5
Opus 43.5
gpt-oss-120b3.6
DeepSeek-R19.3
DeepSeek-V3.111.4
DeepSeek-R1-052825.7
Loading Atlas data…