Thematic Generalization V2 Inverse-Rank Score
General QA · 2026-03-16
Auxiliary Thematic Generalization V2 metric that rewards ranking the hidden theme highly even when it is not the top answer. Higher scores indicate better latent-theme induction.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.6 | 80.6 |
| GPT-5.4 | 80.0 |
| Gemini 3.1 Pro Preview | 79.4 |
| Sonnet 4.6 | 76.3 |
| Opus 4.7 | 72.8 |
| GLM-5.1 | 69.8 |
| Kimi K2.5 | 69.4 |
| Qwen3.5 397B A17B | 65.1 |
| DeepSeek-V3.2 | 65.0 |
| Grok 4.20 | 63.8 |
| Gemini 3.1 Flash-Lite Preview | 63.3 |
| GPT-5.4 Mini | 61.7 |
| Qwen3.6 Plus (2026-04-02) | 59.5 |
| Doubao Seed 2.0 Pro | 57.1 |
| Gemma 4 31B IT | 53.0 |
| Qwen3.5 122B A10B | 51.2 |
| MiMo-V2-Pro | 45.9 |
| Qwen3.5-27B | 45.5 |
| ERNIE 5.0 | 41.7 |
| Trinity Large Thinking | 41.6 |
| MiniMax M2.7 | 39.3 |
| Mistral Large 3 675B Instruct 2512 | 23.0 |
| Mistral Medium 3.1 | 20.3 |