Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Conditional Total Inconsistency

Safety · 2026-03-08

Conditional total inconsistency measures inconsistency among decisive answers in the LLM Sycophancy Benchmark. Lower values indicate more consistent behavior.

Top models (lower is better)

ModelScore
Grok 4.35.3
Grok 4.205.4
Gemini 3.5 Flash11.8
MiniMax M313.8
ERNIE 5.114.7
DeepSeek-V3.217.0
GPT-5.517.1
Qwen3.5 397B A17B18.1
DeepSeek-V4-Pro18.7
Gemma 4 31B IT18.7
GLM-5.119.8
ERNIE 5.019.8
Hy3 preview20.2
Sonnet 4.621.4
MiMo-V2.5-Pro21.7
Gemini 3.1 Flash-Lite Preview22.1
Opus 4.622.9
GLM-523.2
Qwen3.7-Plus23.3
Opus 4.724.6
Kimi K2.625.0
Kimi K2.526.2
Qwen3.6-Max-Preview26.4
Claude Fable 526.6
MiniMax M2.727.4
Doubao Seed 2.0 Pro28.2
Gemini 3.1 Pro Preview28.5
MiniMax M2.530.1
Trinity Large Thinking31.0
Mistral Medium 3.532.7
GPT-4.135.5
Mistral Large 3 675B Instruct 251255.9
Loading Atlas data…