Atlas

Benchmarks

← All benchmarks

PostTrainBench v1.2

Indexes · 2026-09-30

Agents post-train four base models with ten hours on one GPU. Version 1.2 removes BFCL, reweights six downstream benchmarks by the inverse instruct-versus-base gap, fixes HumanEval sandbox scoring, and averages final evaluations over five AIME seeds, three GSM8K/GPQA/HumanEval seeds and one Arena Hard/HealthBench seed. Contamination decisions use a majority of three judge runs. Mixed-agent fallback aggregates are excluded.

Top models (higher is better)

ModelScore
Claude Opus 5.543.8
GPT-6 Astra41.9
Claude Opus 541.7
GPT-6.1 Sol38.8
GLM-5.3 Flash36.8
GLM-5.334.5
GPT-5.6 Sol31.0
Opus 4.831.0
GLM-5.230.4
Opus 4.724.3
GPT-5.523.9
Grok 4.522.6
Gemini 3.1 Pro Preview18.1
GPT-5.417.8
Loading Atlas data…