PostTrainBench v1.2
Indexes · 2026-09-30
Agents post-train four base models with ten hours on one GPU. Version 1.2 removes BFCL, reweights six downstream benchmarks by the inverse instruct-versus-base gap, fixes HumanEval sandbox scoring, and averages final evaluations over five AIME seeds, three GSM8K/GPQA/HumanEval seeds and one Arena Hard/HealthBench seed. Contamination decisions use a majority of three judge runs. Mixed-agent fallback aggregates are excluded.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 43.8 |
| GPT-6 Astra | 41.9 |
| Claude Opus 5 | 41.7 |
| GPT-6.1 Sol | 38.8 |
| GLM-5.3 Flash | 36.8 |
| GLM-5.3 | 34.5 |
| GPT-5.6 Sol | 31.0 |
| Opus 4.8 | 31.0 |
| GLM-5.2 | 30.4 |
| Opus 4.7 | 24.3 |
| GPT-5.5 | 23.9 |
| Grok 4.5 | 22.6 |
| Gemini 3.1 Pro Preview | 18.1 |
| GPT-5.4 | 17.8 |