Atlas

Benchmarks

← All benchmarks

MysteryMechanism

Science · 2026-09-23

Functional recovery of 222 sealed mathematical mechanisms. Agents receive anonymous bounded inputs, two noisy observations, a persistent shell and 2d+1 experiments for d dimensions, without internet or domain context. One submitted executable law is graded on private structural probes using noise-aware error thresholds; provider and agent failures score zero.

Top models (higher is better)

ModelScore
GPT-6 Astra53.2
Claude Opus 5.549.5
Claude Fable 5.147.7
Claude Opus 537.4
Gemini 3.8 Flash36.5
Muse Spark 1.336.0
GPT-5.6 Sol33.3
Grok 4.630.6
Grok 4.725.2
GLM-5.323.0
MiMo-V2.6-Flash21.6
DeepSeek V4.1 Flash21.2
GPT-6 Luna19.4
MiMo-V2.6-Pro15.3
GPT-5.6 Luna14.4
Loading Atlas data…