OEIS Open Lite ($200)
Math · 2026-08-12
Random 100-conjecture OEIS Open subset evaluated with a $200 budget per conjecture using the basic Inspect ReAct agent and SafeVerify. This larger-budget subset is a distinct evaluation protocol from the full $50 benchmark.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 44.0 |
| GPT-5.6 Sol | 43.0 |
| Opus 4.8 | 39.0 |
| GPT-5.5 | 36.0 |
| Gemini 3.5 Flash | 29.0 |