Atlas

Benchmarks

← All benchmarks

Concept Synth INDUCTION v1 - CI

General QA · 2026-05-23

The contrastive-induction split of INDUCTION v1. A formula must exactly match targets in YES worlds while deliberately failing to match at least one target in each contrastive NO world.

Top models (higher is better)

ModelScore
GPT-5.282.5
Opus 4.679.5
GPT-5.479.0
Gemini 3.5 Flash78.5
Grok 478.0
Gemini 3.1 Pro Preview70.0
DeepSeek-V4-Pro67.0
Grok 4.1 Fast60.5
Gemini 3 Pro Preview55.0
Grok 4.346.0
DeepSeek-V3.241.7
Opus 4.534.0
GPT-4o0.5
Loading Atlas data…