Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Balanced 198 Stripped Rate

Safety · 2026-09-30

FIRST on both opposing stripped first-person views in the balanced 198-case cohort. Emotional framing is removed while narrator perspective is retained. Complete 990-prompt runs only.

Top models (lower is better)

ModelScore
Grok 4.50.0
GPT-5.6 Luna0.0
Gemini 3.5 Flash-Lite0.5
Hy31.0
GPT-5.6 Terra1.5
Gemini 3.6 Flash2.5
GPT-5.6 Sol2.5
Kimi K35.1
Inkling5.6
Qwen3.7 Flash6.6
Sonnet 58.1
GLM-5.213.6
Doubao Seed 2.1 Pro17.7
Loading Atlas data…