KernelBench Mega Normalized (B200)
Code · 2026-09-30
KernelBench Mega B200 board: best audited Kimi linear-decode cell relative to the board-best valid cell, scaled to 0–100. Failed/invalid cells contribute zero. The retired RL-grid problem is absent. Best published cells may span harness routes and reasoning settings; distinct from absolute speedup.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 100.0 |
| Opus 4.8 | 54.6 |
| GPT-5.5 | 26.4 |
| GLM-5.2 | 20.6 |
| MiniMax M3 | 11.7 |
| Gemini 3.5 Flash | 7.2 |
| Composer 2.5 | 3.3 |
| Grok 4.5 | 0.1 |
| Kimi K2.7 Code | 0.0 |