Vals RSI Index v1.1
Agents · 2026-09-21
Mean of five autonomous AI research campaign scores on reference-anchored logarithmic scales. Zero is the starting baseline, 50 a selected published, human or model reference, and 100 the theoretical best. Each campaign uses a native coding agent, maximum reported effort, a persistent Marimo notebook and a fixed 12–30 hour resource budget. References differ across tasks and do not uniformly represent human ability.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 37.1 |
| Claude Fable 5.1 | 35.0 |
| Claude Opus 5 | 32.1 |
| GPT-6 Astra | 29.0 |
| Claude Fable 5 | 24.5 |
| Opus 4.8 | 23.2 |
| GLM-5.3 | 23.1 |
| GPT-5.6 Sol | 23.1 |
| Kimi K3 | 21.7 |
| Grok 4.6 | 21.5 |
| Opus 4.7 | 20.1 |
| Muse Spark 1.3 | 18.7 |
| GPT-5.5 | 17.9 |
| Gemini 3.8 Flash | 17.6 |
| GPT-5.4 | 15.3 |
| GPT-5.2 | 15.2 |
| Kimi K2.5 | 5.2 |