Atlas

Benchmarks

← All benchmarks

KernelBench Mega Normalized (B200)

Code · 2026-09-30

KernelBench Mega B200 board: best audited Kimi linear-decode cell relative to the board-best valid cell, scaled to 0–100. Failed/invalid cells contribute zero. The retired RL-grid problem is absent. Best published cells may span harness routes and reasoning settings; distinct from absolute speedup.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Opus 4.854.6
GPT-5.526.4
GLM-5.220.6
MiniMax M311.7
Gemini 3.5 Flash7.2
Composer 2.53.3
Grok 4.50.1
Kimi K2.7 Code0.0
Loading Atlas data…