Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Conditional Rate

Safety · 2026-03-08

Conditional sycophancy rate measures sycophantic behavior among questions where the model gives decisive answers. Lower values indicate less sycophancy.

Top models (lower is better)

ModelScore
Claude Fable 50.6
Gemini 3.1 Pro Preview0.7
Qwen3.5 397B A17B3.4
Grok 4.203.6
Opus 4.64.2
GPT-5.54.6
Grok 4.35.3
MiMo-V2.5-Pro5.4
Qwen3.6-Max-Preview5.6
Gemini 3.5 Flash5.9
ERNIE 5.15.9
Gemini 3.1 Flash-Lite Preview6.3
Gemma 4 31B IT6.7
Opus 4.76.9
Kimi K2.68.3
GLM-5.19.4
Qwen3.7-Plus9.7
Sonnet 4.69.7
MiniMax M310.8
DeepSeek-V4-Pro11.0
DeepSeek-V3.211.3
MiniMax M2.511.8
Kimi K2.512.1
GLM-513.0
MiniMax M2.715.3
Hy3 preview15.4
ERNIE 5.016.7
Doubao Seed 2.0 Pro25.5
Trinity Large Thinking25.5
Mistral Medium 3.526.7
GPT-4.134.5
Mistral Large 3 675B Instruct 251252.5
Loading Atlas data…