Atlas

Benchmarks

← All benchmarks

Thematic Generalization V2 Inverse-Rank Score

General QA · 2026-03-16

Auxiliary Thematic Generalization V2 metric that rewards ranking the hidden theme highly even when it is not the top answer. Higher scores indicate better latent-theme induction.

Top models (higher is better)

ModelScore
Opus 4.680.6
GPT-5.480.0
Gemini 3.1 Pro Preview79.4
Sonnet 4.676.3
Opus 4.772.8
GLM-5.169.8
Kimi K2.569.4
Qwen3.5 397B A17B65.1
DeepSeek-V3.265.0
Grok 4.2063.8
Gemini 3.1 Flash-Lite Preview63.3
GPT-5.4 Mini61.7
Qwen3.6 Plus (2026-04-02)59.5
Doubao Seed 2.0 Pro57.1
Gemma 4 31B IT53.0
Qwen3.5 122B A10B51.2
MiMo-V2-Pro45.9
Qwen3.5-27B45.5
ERNIE 5.041.7
Trinity Large Thinking41.6
MiniMax M2.739.3
Mistral Large 3 675B Instruct 251223.0
Mistral Medium 3.120.3
Loading Atlas data…