Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Rate

Safety · 2026-03-08

Sycophancy rate is the share of decisive preference questions on which the model shows sycophantic behavior. Lower values indicate less sycophancy.

Top models (lower is better)

ModelScore
Claude Fable 50.5
Gemini 3.1 Pro Preview0.5
Grok 4.30.5
Gemini 3.5 Flash0.5
Grok 4.201.0
Qwen3.5 397B A17B2.0
ERNIE 5.12.0
Opus 4.62.5
MiMo-V2.5-Pro2.5
Gemini 3.1 Flash-Lite Preview3.0
GPT-5.53.5
MiniMax M33.5
Qwen3.6-Max-Preview4.0
Gemma 4 31B IT4.5
Opus 4.74.5
Kimi K2.64.5
GLM-5.14.5
Qwen3.7-Plus5.0
DeepSeek-V4-Pro5.0
DeepSeek-V3.26.0
Kimi K2.56.5
Sonnet 4.67.0
Hy3 preview8.0
ERNIE 5.08.0
MiniMax M2.59.0
MiniMax M2.79.5
GLM-512.1
Doubao Seed 2.0 Pro14.1
Trinity Large Thinking18.6
GPT-4.119.1
Mistral Medium 3.522.1
Mistral Large 3 675B Instruct 251231.2
Loading Atlas data…