Concept Synth ABD v1.2
General QA · 2026-02-21
ABD v1.2 tests whether a model can synthesize one first-order abnormality rule that repairs a default theory across multiple finite relational worlds. The official headline score is prompt validity after the release's conservative suffix repair, checked exactly with SMT across 600 instances in three observation regimes.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.6 | 98.8 |
| Gemini 3.1 Pro Preview | 98.0 |
| DeepSeek-V3.2 | 95.9 |
| Grok 4.1 Fast | 95.2 |
| GPT-5.2 | 92.2 |
| Grok 4 | 89.4 |
| GPT-5.4 | 85.2 |
| Gemini 3 Pro Preview | 76.9 |
| Kimi K2 Thinking | 71.5 |
| GPT-4o | 19.8 |