Atlas

Benchmarks

← All benchmarks

Vals RSI (four tasks) — Harness Engineering

Agents

Reference-normalized judge-harness research after 12 hours. Agents improve a fixed Gemini 3.5 Flash-Lite judge using 12 development apps; the frozen harness is evaluated three times on 24 held-out apps with 205 substeps. Agreement baseline 50%, reference 93.6%, theoretical best 99%. The changed reference separates this panel from the earlier five-task RSI protocol.

Top models (higher is better)

ModelScore
GPT-6 Sol22.8
Claude Fable 5.122.6
Grok 4.621.8
GPT-5.6 Sol18.5
Claude Fable 518.1
GLM-5.317.4
Opus 4.816.8
Claude Opus 514.8
Claude Opus 5.514.5
Gemini 3.8 Flash14.5
Muse Spark 1.314.5
Grok 4.711.3
Kimi K39.9
GPT-6 Astra9.5
GPT-5.48.7
Gemini 3.7 Flash8.4
Opus 4.75.6
GPT-5.55.6
GPT-5.23.9
Kimi K2.52.1
Loading Atlas data…