Atlas

Benchmarks

← All benchmarks

VoxelBench Text Rating

Games · 2025-09-12

VoxelBench text rating is the Glicko-2 leaderboard track for models creating voxel builds from text prompts. Ratings come from head-to-head human votes over rendered voxel constructions.

Top models (higher is better)

ModelScore
GPT-5.6 Sol2291
Claude Opus 52276
Claude Fable 52225
Kimi K32015
GPT-5.5 Pro1997
GPT-5.51980
Grok 4.51803
Gemini 3.6 Flash1787
Opus 4.81681
Gemini 3.1 Pro Preview1681
Gemini 3.5 Flash1675
Qwen3.7-Max1625
Sonnet 51595
GLM-5.21591
Muse Spark 1.11518
GPT-5.41513
Opus 4.51504
Gemini 2.5 Deep Think1491
Opus 4.61478
GPT-5.21447
Gemini 3 Deep Think1446
GLM-5.11420
MiniMax M31420
Opus 4.71415
Gemini 3 Pro Preview1359
Gemini 3 Flash Preview1316
Qwen3.6 Plus (2026-04-02)1308
GPT-51300
Gemini 2.5 Pro1257
GLM-51252
GPT-5.4 Mini1251
GPT-5-Codex1251
Kimi K2.61243
GPT-5 Pro1233
DeepSeek-V4-Pro1227
Sonnet 4.51226
Opus 4.11223
GPT-5.11221
Sonnet 41138
DeepSeek-V4-Flash1133
Loading Atlas data…