KernelBench Hard Normalized (B200)
Code · 2026-09-30
KernelBench Hard B200 board: mean audited valid cell score relative to the board-best score per problem, scaled to 0–100 over six problems. Missing/invalid cells contribute zero. Best published cells may come from different harness routes and reasoning settings of the same model.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 94.3 |
| Kimi K3 | 77.0 |
| GPT-5.6 Sol | 68.9 |
| Opus 4.8 | 10.5 |
| Grok 4.5 | 9.3 |
| Composer 2.5 | 0.0 |
| Gemini 3.5 Flash | 0.0 |
| GLM-5.2 | 0.0 |
| GPT-5.5 | 0.0 |
| Kimi K2.7 Code | 0.0 |
| MiniMax M3 | 0.0 |