Atlas

Benchmarks

← All benchmarks

Concept Synth INDUCTION v1 - EC

General QA · 2026-05-23

The existential-completion split of INDUCTION v1. A formula is valid when every partially observed input world has some completion under which the formula matches all target labels, checked exactly with Z3.

Top models (higher is better)

ModelScore
GPT-5.494.0
GPT-5.278.0
Opus 4.658.0
DeepSeek-V4-Pro53.8
Gemini 3 Pro Preview53.5
Grok 453.3
Gemini 3.1 Pro Preview53.0
Gemini 3.5 Flash51.5
Grok 4.1 Fast41.0
Grok 4.338.5
DeepSeek-V3.233.0
Opus 4.530.0
GPT-4o2.0
Loading Atlas data…