LLM Sycophancy Balanced 198 Insufficient Rate
Safety · 2026-09-30
Percentage of individual responses labeled INSUFFICIENT among all 990 prompts in the balanced 198-case cohort. It uses individual prompt responses, not case pairs, as its denominator.
Top models (lower is better)
| Model | Score |
|---|---|
| Kimi K3 | 4.7 |
| Inkling | 13.5 |
| Doubao Seed 2.1 Pro | 13.6 |
| GLM-5.2 | 17.7 |
| Sonnet 5 | 22.5 |
| GPT-5.6 Sol | 22.6 |
| Qwen3.7 Flash | 48.4 |
| GPT-5.6 Luna | 49.4 |
| GPT-5.6 Terra | 49.9 |
| Gemini 3.6 Flash | 60.7 |
| Grok 4.5 | 64.1 |
| Hy3 | 68.5 |
| Gemini 3.5 Flash-Lite | 83.9 |