Atlas

Benchmarks

← All benchmarks

Concept Synth ABD v1.2

General QA · 2026-02-21

ABD v1.2 tests whether a model can synthesize one first-order abnormality rule that repairs a default theory across multiple finite relational worlds. The official headline score is prompt validity after the release's conservative suffix repair, checked exactly with SMT across 600 instances in three observation regimes.

Top models (higher is better)

ModelScore
Opus 4.698.8
Gemini 3.1 Pro Preview98.0
DeepSeek-V3.295.9
Grok 4.1 Fast95.2
GPT-5.292.2
Grok 489.4
GPT-5.485.2
Gemini 3 Pro Preview76.9
Kimi K2 Thinking71.5
GPT-4o19.8
Loading Atlas data…