Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Insufficient Answer Rate

Safety · 2026-03-08

Insufficient answer rate is the percentage of sycophancy-test items where the model does not provide a decisive answer. Lower values indicate fewer insufficient responses.

Top models (lower is better)

ModelScore
GLM-53.9
Mistral Medium 3.510.5
Trinity Large Thinking14.1
MiniMax M2.514.8
GPT-5.516.8
Sonnet 4.617.6
Qwen3.6-Max-Preview21.9
MiniMax M2.723.2
Mistral Large 3 675B Instruct 251226.1
Gemma 4 31B IT27.6
Claude Fable 527.7
Opus 4.627.7
Gemini 3.1 Pro Preview28.2
Opus 4.728.6
Qwen3.5 397B A17B29.9
Doubao Seed 2.0 Pro31.1
GPT-4.131.2
Kimi K2.631.6
Kimi K2.533.5
Hy3 preview34.3
GLM-5.134.8
Qwen3.7-Plus35.9
DeepSeek-V3.236.8
ERNIE 5.037.2
MiMo-V2.5-Pro37.4
DeepSeek-V4-Pro38.4
Gemini 3.1 Flash-Lite Preview43.5
ERNIE 5.143.7
MiniMax M348.7
Grok 4.2060.9
Grok 4.380.2
Gemini 3.5 Flash87.1
Loading Atlas data…