Atlas

Benchmarks

← All benchmarks

Vals RSI — Parameter Golf

Agents · 2026-09-21

Reference-normalized score for one 12-hour offline research campaign on eight H100s. Final language-model artifact must fit 16 MB and train in ten minutes; rule review and three hidden-seed retrainings determine BPB. Baseline 1.2288 BPB, reference 1.0565, theoretical best 0.6.

Top models (higher is better)

ModelScore
Claude Opus 5.541.4
GPT-6 Astra40.6
Claude Fable 5.137.0
Claude Opus 533.4
Claude Fable 530.6
GLM-5.329.5
Kimi K329.0
Opus 4.726.7
GPT-5.526.1
Opus 4.825.6
GPT-5.6 Sol25.6
Muse Spark 1.319.8
GPT-5.418.1
GPT-5.214.9
Gemini 3.8 Flash13.1
Grok 4.611.9
Kimi K2.50.0
Loading Atlas data…