Vals RSI — Post-training
Agents · 2026-09-21
Reference-normalized score for improving a provided Qwen3.6-35B-A3B checkpoint in 30 hours on sixteen RTX PRO 6000 GPUs, with internet, 27 public Finance Agent v2 examples and $100 external inference credits. One submitted checkpoint and generation setup is evaluated three times on 126 held-out tasks. Baseline 35.38%, reference 59.42%, theoretical best 99%.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 35.1 |
| Claude Opus 5 | 31.8 |
| Claude Fable 5.1 | 28.0 |
| GPT-6 Astra | 7.8 |
| Claude Fable 5 | 0.0 |
| Opus 4.7 | 0.0 |
| Opus 4.8 | 0.0 |
| Gemini 3.8 Flash | 0.0 |
| GLM-5.3 | 0.0 |
| GPT-5.2 | 0.0 |
| GPT-5.4 | 0.0 |
| GPT-5.5 | 0.0 |
| GPT-5.6 Sol | 0.0 |
| Grok 4.6 | 0.0 |
| Kimi K2.5 | 0.0 |
| Kimi K3 | 0.0 |
| Muse Spark 1.3 | 0.0 |