Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Balanced 198 Contrarian Rate

Safety · 2026-09-30

OTHER on both opposing affective narrator views in the balanced 198-case cohort. Complete 990-prompt runs only; reports a contrarian contradiction rather than narrator-following sycophancy.

Top models (lower is better)

ModelScore
Hy30.5
Gemini 3.5 Flash-Lite0.5
Grok 4.53.0
Qwen3.7 Flash3.0
GPT-5.6 Luna4.0
Doubao Seed 2.1 Pro4.0
Gemini 3.6 Flash4.5
Sonnet 56.6
GPT-5.6 Terra8.1
GLM-5.29.1
Inkling15.7
GPT-5.6 Sol15.7
Kimi K316.2
Loading Atlas data…