LLM Position Bias Decisive Pair Coverage
Chat & Writing · 2026-04-21
Decisive Pair Coverage is the share of cases where the model picked one side in both swapped views. The benchmark uses paired story comparisons under order swaps.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.7 | 94.8 |
| Grok 4.20 | 94.3 |
| Kimi K2.6 | 94.3 |
| Kimi K2.5 | 94.3 |
| Qwen3.7-Max | 93.8 |
| GPT-5.4 | 93.8 |
| GPT-5.5 | 93.3 |
| Opus 4.8 | 92.7 |
| Sonnet 4.6 | 92.7 |
| Gemini 3.5 Flash | 92.2 |
| Gemini 3.1 Pro Preview | 92.2 |
| Qwen3.5 122B A10B | 90.2 |
| Gemma 4 31B IT | 89.6 |
| DeepSeek-V4-Pro | 89.1 |
| Mistral Large 3 675B Instruct 2512 | 89.1 |
| Qwen3.5 397B A17B | 88.6 |
| GPT-5.4 Mini | 88.1 |
| GLM-5.1 | 87.0 |
| Mistral Medium 3.1 | 85.0 |
| Doubao Seed 2.0 Pro | 83.4 |
| DeepSeek-V3.2 | 75.1 |
| Trinity Large Thinking | 75.1 |
| Qwen3.6 Plus (2026-04-02) | 73.6 |
| ERNIE 5.0 | 72.5 |
| Claude Fable 5 | 72.0 |
| MiniMax M2.7 | 65.3 |
| Opus 4.6 | 60.1 |
| Mistral Medium 3.5 | 56.5 |
| MiMo-V2-Pro | 54.9 |
| Gemini 3.1 Flash-Lite Preview | 54.9 |
| MiniMax M3 | 44.6 |