Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Stripped Rate

Safety · 2026-03-08

Stripped sycophancy rate measures sycophancy after removing user-opinion cues from prompts. Lower values indicate less sycophancy.

Top models (lower is better)

ModelScore
Gemini 3.1 Pro Preview0.0
Grok 4.30.0
Gemini 3.5 Flash0.0
Gemini 3.1 Flash-Lite Preview0.5
Claude Fable 52.0
Grok 4.202.0
Opus 4.72.0
MiMo-V2.5-Pro2.5
Kimi K2.63.0
Kimi K2.53.0
GPT-5.53.5
GLM-5.13.5
Qwen3.5 397B A17B4.0
DeepSeek-V4-Pro4.0
DeepSeek-V3.24.0
Qwen3.6-Max-Preview4.5
Gemma 4 31B IT5.0
Qwen3.7-Plus5.0
Opus 4.66.0
MiniMax M2.56.0
MiniMax M37.0
ERNIE 5.17.5
Sonnet 4.67.5
ERNIE 5.08.0
Hy3 preview10.1
GLM-510.1
MiniMax M2.714.1
Trinity Large Thinking15.1
Doubao Seed 2.0 Pro17.1
GPT-4.118.1
Mistral Medium 3.524.1
Mistral Large 3 675B Instruct 251231.2
Loading Atlas data…