Vals RSI Index v1.1 (four tasks)
Agents
Mean of four autonomous AI research campaign scores on reference-anchored logarithmic scales, first captured September 28, 2026. This revision removes Parameter Golf and changes the harness-engineering reference to 93.6% agreement. Zero is the starting baseline, 50 a selected published, human or model reference, and 100 the theoretical best. Native coding agents run with maximum reported effort, persistent Marimo notebooks and fixed 12–30 hour budgets. This protocol is distinct from the earlier five-task index.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 37.3 |
| Claude Fable 5.1 | 36.1 |
| Claude Opus 5 | 33.0 |
| GPT-6 Sol | 28.1 |
| GPT-6 Astra | 27.1 |
| Grok 4.6 | 25.5 |
| Claude Fable 5 | 24.3 |
| Opus 4.8 | 24.0 |
| GPT-5.6 Sol | 23.9 |
| GLM-5.3 | 22.9 |
| Kimi K3 | 20.9 |
| Grok 4.7 | 20.2 |
| Gemini 3.8 Flash | 19.9 |
| Muse Spark 1.3 | 19.6 |
| Opus 4.7 | 19.1 |
| Gemini 3.7 Flash | 18.3 |
| GPT-5.5 | 16.5 |
| GPT-5.2 | 15.8 |
| GPT-5.4 | 15.5 |
| Kimi K2.5 | 6.8 |