Atlas

Benchmarks

← All benchmarks

Vals RSI Index v1.1 (four tasks)

Agents

Mean of four autonomous AI research campaign scores on reference-anchored logarithmic scales, first captured September 28, 2026. This revision removes Parameter Golf and changes the harness-engineering reference to 93.6% agreement. Zero is the starting baseline, 50 a selected published, human or model reference, and 100 the theoretical best. Native coding agents run with maximum reported effort, persistent Marimo notebooks and fixed 12–30 hour budgets. This protocol is distinct from the earlier five-task index.

Top models (higher is better)

ModelScore
Claude Opus 5.537.3
Claude Fable 5.136.1
Claude Opus 533.0
GPT-6 Sol28.1
GPT-6 Astra27.1
Grok 4.625.5
Claude Fable 524.3
Opus 4.824.0
GPT-5.6 Sol23.9
GLM-5.322.9
Kimi K320.9
Grok 4.720.2
Gemini 3.8 Flash19.9
Muse Spark 1.319.6
Opus 4.719.1
Gemini 3.7 Flash18.3
GPT-5.516.5
GPT-5.215.8
GPT-5.415.5
Kimi K2.56.8
Loading Atlas data…