Atlas

Benchmarks

← All benchmarks

Concept Synth ABD v1.2 - Skeptical

General QA · 2026-02-21

The partially observed, robust-completion split of ABD v1.2. A rule is valid only when the repaired theory is satisfied under every completion of the unknown atoms in every prompt world.

Top models (higher is better)

ModelScore
GPT-5.2100.0
Opus 4.698.8
Gemini 3.1 Pro Preview97.5
DeepSeek-V3.295.7
Grok 4.1 Fast95.5
GPT-5.493.8
Gemini 3 Pro Preview92.6
Grok 488.9
Kimi K2 Thinking82.7
GPT-4o29.0
Loading Atlas data…