Atlas

Benchmarks

← All benchmarks

LLM Sycophancy Decisive Coverage

Safety · 2026-03-08

Decisive coverage is the percentage of sycophancy-test items for which the model gives a decisive answer. Higher values indicate greater coverage of the test questions, though coverage should be interpreted alongside sycophancy rates.

Top models (higher is better)

ModelScore
GLM-593.0
Mistral Medium 3.582.9
Claude Fable 577.4
MiniMax M2.576.9
GPT-5.576.4
Gemini 3.1 Pro Preview75.9
Sonnet 4.672.9
Trinity Large Thinking72.9
Qwen3.6-Max-Preview72.4
Gemma 4 31B IT67.3
Opus 4.765.3
MiniMax M2.762.3
Opus 4.659.3
Mistral Large 3 675B Instruct 251259.3
Qwen3.5 397B A17B58.3
Doubao Seed 2.0 Pro55.3
GPT-4.155.3
Kimi K2.654.3
Kimi K2.553.8
DeepSeek-V3.253.3
Hy3 preview52.3
Qwen3.7-Plus51.8
GLM-5.148.2
ERNIE 5.048.2
Gemini 3.1 Flash-Lite Preview47.7
MiMo-V2.5-Pro46.2
DeepSeek-V4-Pro45.7
ERNIE 5.134.2
MiniMax M332.7
Grok 4.2028.1
Grok 4.39.5
Gemini 3.5 Flash8.5
Loading Atlas data…