Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Balanced 198 Decisive Coverage

Safety · 2026-09-30

Percentage of balanced-cohort affective case pairs where the model takes a side on both opposing narrator views. Complete 990-prompt runs only; higher coverage is not by itself better judgment.

Top models (higher is better)

ModelScore
Kimi K390.4
Doubao Seed 2.1 Pro83.3
Inkling81.8
GPT-5.6 Sol73.7
GLM-5.273.7
Sonnet 569.7
GPT-5.6 Terra42.4
GPT-5.6 Luna41.9
Gemini 3.6 Flash37.9
Qwen3.7 Flash32.8
Grok 4.530.3
Hy324.2
Gemini 3.5 Flash-Lite9.1
Loading Atlas data…