Atlas

Benchmarks

← All benchmarks

Thematic Generalization V2

General QA · 2026-03-16

Thematic Generalization V2 Top-1 Accuracy on the pinned 703-case hard subset tests whether models can infer a latent theme from positive examples and anti-examples and identify the one true match among close distractors. Higher is better; inverse-rank results use the separate thematic-generalization-v2-inverse-rank-score slug.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview91.5
GPT-5.490.9
Opus 4.690.0
Sonnet 4.688.5
Opus 4.786.8
GLM-5.185.9
Kimi K2.584.8
Qwen3.5 397B A17B82.4
Gemini 3.1 Flash-Lite Preview82.1
DeepSeek-V3.281.8
Grok 4.2081.5
Qwen3.6 Plus (2026-04-02)81.5
GPT-5.4 Mini80.8
Doubao Seed 2.0 Pro77.0
Qwen3.5 122B A10B76.5
Gemma 4 31B IT76.1
Qwen3.5-27B71.3
MiMo-V2-Pro68.8
Trinity Large Thinking66.1
ERNIE 5.065.3
MiniMax M2.763.4
Mistral Large 3 675B Instruct 251234.0
Mistral Medium 3.132.3
Loading Atlas data…