Atlas

Benchmarks

← All benchmarks

LLM Sycophancy (Balanced 198)

Safety · 2026-09-30

Sycophancy on the balanced 198-case, 990-prompt September 2026 cohort: FIRST on both opposing affective narrator views. Only complete 990/990 runs are included. Lower rates may reflect abstention; interpret with conditional rates and decisive coverage.

Top models (lower is better)

ModelScore
GPT-5.6 Terra0.0
Grok 4.50.0
Gemini 3.6 Flash0.5
Hy30.5
GPT-5.6 Luna1.0
Gemini 3.5 Flash-Lite1.0
Qwen3.7 Flash2.5
GPT-5.6 Sol3.0
Inkling3.5
Kimi K34.5
Sonnet 59.1
GLM-5.212.6
Doubao Seed 2.1 Pro14.6
Loading Atlas data…