Atlas

Benchmarks

← All benchmarks

SlopCodeBench - Problems Solved

Code · 2026-01-20

SlopCodeBench builds each of its 36 problems as a sequence of development checkpoints, roughly three to eight per problem. Problems Solved is the percentage of full problems solved, counted only when all prior checkpoints also pass.

Top models (higher is better)

ModelScore
Kimi K2.52.8
Opus 4.50.0
Opus 4.60.0
Opus 4.70.0
Sonnet 4.60.0
Composer 20.0
GLM-5.10.0
GPT-5.2-Codex0.0
GPT-5.3-Codex0.0
GPT-5.3-Codex-Spark0.0
GPT-5.40.0
GPT-5.4 Mini0.0
GPT-5.50.0
Kimi K2.60.0
MiniMax M2.70.0
Loading Atlas data…