Atlas

Benchmarks

← All benchmarks

OEIS Open Lite ($200)

Math · 2026-08-12

Random 100-conjecture OEIS Open subset evaluated with a $200 budget per conjecture using the basic Inspect ReAct agent and SafeVerify. This larger-budget subset is a distinct evaluation protocol from the full $50 benchmark.

Top models (higher is better)

ModelScore
Claude Fable 544.0
GPT-5.6 Sol43.0
Opus 4.839.0
GPT-5.536.0
Gemini 3.5 Flash29.0
Loading Atlas data…