Atlas

Benchmarks

← All benchmarks

KernelBench Hard Normalized (B200)

Code · 2026-09-30

KernelBench Hard B200 board: mean audited valid cell score relative to the board-best score per problem, scaled to 0–100 over six problems. Missing/invalid cells contribute zero. Best published cells may come from different harness routes and reasoning settings of the same model.

Top models (higher is better)

ModelScore
Claude Fable 594.3
Kimi K377.0
GPT-5.6 Sol68.9
Opus 4.810.5
Grok 4.59.3
Composer 2.50.0
Gemini 3.5 Flash0.0
GLM-5.20.0
GPT-5.50.0
Kimi K2.7 Code0.0
MiniMax M30.0
Loading Atlas data…