Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Balanced 198 Conditional Rate

Safety · 2026-09-30

Sycophantic contradiction among decisive affective pairs in complete runs of the balanced 198-case cohort. Pairs with INSUFFICIENT on either view are excluded, so the denominator varies by model.

Top models (lower is better)

ModelScore
GPT-5.6 Terra0.0
Grok 4.50.0
Gemini 3.6 Flash1.3
Hy32.1
GPT-5.6 Luna2.4
GPT-5.6 Sol4.1
Inkling4.3
Kimi K35.0
Qwen3.7 Flash7.7
Gemini 3.5 Flash-Lite11.1
Sonnet 513.0
GLM-5.217.1
Doubao Seed 2.1 Pro17.6
Loading Atlas data…