Atlas

Benchmarks

← All benchmarks

BLXBench UI category

Code · 2026-05-10

UI category component of the BLXBench public leaderboard. The source describes this category as single-file HTML visual/UI artifacts with render and preview workflows; scores are category score percentages.

Top models (higher is better)

ModelScore
GPT-5.587.5
Opus 4.787.0
Claude Fable 586.2
Qwen3.7-Max86.0
Grok 4.385.6
MiMo-V2.585.2
Opus 4.884.5
DeepSeek-V4-Pro82.6
DeepSeek-V4-Flash79.1
Qwen3.7-Plus78.9
Mistral Small 477.1
Kimi K2.676.9
Mistral Medium 3.576.9
MiMo-V2.5-Pro75.1
GPT-5.3-Codex68.0
Grok Build 0.165.4
Kimi K2.7 Code63.0
Qwen3.6-Flash-2026-04-1662.5
CoBuddy61.1
MiniMax M355.8
North Mini Code51.5
Nemotron 3 Super 120B A12B43.2
Ring 2.6 1T35.4
Gemini 3.1 Flash-Lite34.8
Nemotron 3 Nano 30B A3B33.1
GLM-5.130.1
GLM-5.223.7
Granite 4.1 8B21.9
Gemini 3.5 Flash18.8
Nemotron 3 Nano Omni 30B A3B17.6
MiniMax M2.711.9
Step 3.7 Flash6.1
Loading Atlas data…