Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Total Inconsistency

Safety · 2026-03-08

Total inconsistency measures how often the model changes its answer across paired sycophancy-test prompts. Lower values indicate more consistent behavior.

Top models (lower is better)

ModelScore
Grok 4.30.5
Gemini 3.5 Flash1.0
Grok 4.201.5
MiniMax M34.5
ERNIE 5.15.0
DeepSeek-V4-Pro8.5
DeepSeek-V3.29.0
GLM-5.19.5
ERNIE 5.09.5
MiMo-V2.5-Pro10.1
Qwen3.5 397B A17B10.6
Gemini 3.1 Flash-Lite Preview10.6
Hy3 preview10.6
Qwen3.7-Plus12.1
Gemma 4 31B IT12.6
GPT-5.513.1
Opus 4.613.6
Kimi K2.613.6
Kimi K2.514.1
Sonnet 4.615.6
Doubao Seed 2.0 Pro15.6
Opus 4.716.1
MiniMax M2.717.1
Qwen3.6-Max-Preview19.1
GPT-4.119.6
Claude Fable 520.6
Gemini 3.1 Pro Preview21.6
GLM-521.6
Trinity Large Thinking22.6
MiniMax M2.523.1
Mistral Medium 3.527.1
Mistral Large 3 675B Instruct 251233.2
Loading Atlas data…