Atlas

Benchmarks

← All benchmarks

LLM Position Bias Decisive Pair Coverage

Chat & Writing · 2026-04-21

Decisive Pair Coverage is the share of cases where the model picked one side in both swapped views. The benchmark uses paired story comparisons under order swaps.

Top models (higher is better)

ModelScore
Opus 4.794.8
Grok 4.2094.3
Kimi K2.694.3
Kimi K2.594.3
Qwen3.7-Max93.8
GPT-5.493.8
GPT-5.593.3
Opus 4.892.7
Sonnet 4.692.7
Gemini 3.5 Flash92.2
Gemini 3.1 Pro Preview92.2
Qwen3.5 122B A10B90.2
Gemma 4 31B IT89.6
DeepSeek-V4-Pro89.1
Mistral Large 3 675B Instruct 251289.1
Qwen3.5 397B A17B88.6
GPT-5.4 Mini88.1
GLM-5.187.0
Mistral Medium 3.185.0
Doubao Seed 2.0 Pro83.4
DeepSeek-V3.275.1
Trinity Large Thinking75.1
Qwen3.6 Plus (2026-04-02)73.6
ERNIE 5.072.5
Claude Fable 572.0
MiniMax M2.765.3
Opus 4.660.1
Mistral Medium 3.556.5
MiMo-V2-Pro54.9
Gemini 3.1 Flash-Lite Preview54.9
MiniMax M344.6
Loading Atlas data…