Atlas

Benchmarks

← All benchmarks

Concept Synth INDUCTION Challenge64 Round 1

General QA · 2026-07-21

Challenge64 is a 64-instance hard FullObs subset of INDUCTION evaluated on one selected direct Round-1 formula per task before symbolic repair or simplification. Accuracy uses the fixed denominator of 64, so missing or parser-invalid responses count as incorrect; the separately released generated-holdout diagnostic is not included here.

Top models (higher is better)

ModelScore
GPT-5.6 Sol57.8
Claude Fable 540.6
GPT-5.6 Terra37.5
Muse Spark 1.132.8
GPT-5.6 Luna23.4
Grok 420.3
GPT-5.418.8
Grok 4.517.2
GPT-5.214.1
Gemini 3.5 Flash10.9
Kimi K310.9
Grok 4.1 Fast9.4
DeepSeek-V4-Pro9.4
Opus 4.67.8
Gemini 3.6 Flash7.8
Kimi K2.7 Code7.8
Opus 4.86.3
Gemini 3 Pro Preview4.7
Kimi K2.64.7
Sonnet 54.7
Grok 4.33.1
DeepSeek-V3.21.6
Gemini 3.1 Pro Preview1.6
Opus 4.50.0
GPT-4o0.0
Qwen3.7-Max0.0
Loading Atlas data…