PostTrainBench Average
Indexes · 2026-03-09
PostTrainBench evaluates model post-training methods across its task suite. This row reports the benchmark's weighted-average percentage score.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 36.6 |
| GLM-5.2 | 34.3 |
| Opus 4.8 | 34.1 |
| Opus 4.7 | 28.6 |
| GPT-5.5 | 27.2 |
| Grok 4.5 | 23.4 |
| Gemini 3.1 Pro Preview | 22.0 |
| GPT-5.4 | 19.0 |