Atlas

Benchmarks

← All benchmarks

Vals RSI Index v1.1

Agents · 2026-09-21

Mean of five autonomous AI research campaign scores on reference-anchored logarithmic scales. Zero is the starting baseline, 50 a selected published, human or model reference, and 100 the theoretical best. Each campaign uses a native coding agent, maximum reported effort, a persistent Marimo notebook and a fixed 12–30 hour resource budget. References differ across tasks and do not uniformly represent human ability.

Top models (higher is better)

ModelScore
Claude Opus 5.537.1
Claude Fable 5.135.0
Claude Opus 532.1
GPT-6 Astra29.0
Claude Fable 524.5
Opus 4.823.2
GLM-5.323.1
GPT-5.6 Sol23.1
Kimi K321.7
Grok 4.621.5
Opus 4.720.1
Muse Spark 1.318.7
GPT-5.517.9
Gemini 3.8 Flash17.6
GPT-5.415.3
GPT-5.215.2
Kimi K2.55.2
Loading Atlas data…