Atlas

Benchmarks

← All benchmarks

ARC-AGI-2

General QA · 2025-03-24

A harder successor to ARC-AGI that tests few-shot abstract reasoning and pattern generalization on grid-based tasks, with an added emphasis on efficiency of compute per task solved.

Top models (higher is better)

ModelScore
GPT-5.6 Sol92.5
Claude Opus 590.4
GPT-5.585.0
GPT-5.5 Pro84.6
GPT-5.6 Terra83.9
Opus 4.775.8
Gemini 3.5 Flash72.1
Opus 4.871.7
Opus 4.669.2
Sonnet 4.660.4
Kimi K360.4
GPT-5.6 Luna59.5
GPT-5.455.4
Grok 4.552.6
Inkling-Small40.1
GPT-5.2 Pro38.5
Opus 4.537.6
Inkling36.5
Gemini 3 Flash Preview33.6
Gemini 3 Pro Preview31.1
DeepSeek-V4-Pro28.3
GPT-5.226.7
GLM-5.222.8
GPT-5.4 Mini18.9
Kimi K2.618.1
Grok 416.0
GLM-5.115.0
Sonnet 4.513.6
Grok 4.313.3
Kimi K2.511.8
Opus 48.6
GPT-57.5
Grok 4.1 Fast6.7
GPT-5.16.5
GPT-5.4 Nano5.7
Grok 4 Fast5.3
DeepSeek-V3.25.0
Kimi K2 Thinking5.0
Gemini 2.5 Pro4.9
GLM-54.9
Loading Atlas data…