Atlas

Benchmarks

← All benchmarks

MineBench — Bradley-Terry Rating

Games

MineBench ranks Minecraft-style voxel builds from blind human pairwise votes using a global Bradley-Terry model on a 1500-centered Elo scale. Ties contribute half a point; both-bad votes are excluded from the skill fit. This rating replaces the earlier conservative Glicko score and is kept as a separate protocol.

Top models (higher is better)

ModelScore
GPT-6 Astra Pro2262
Claude Opus 5.52150
Claude Opus 52063
GPT-5.6 Sol Pro2057
GPT-6 Sol Pro2057
Claude Fable 5.11984
GPT-5.5 Pro1982
Claude Fable 51925
GPT-5.51904
DeepSeek V4.1 Flash1887
Grok 4.61884
Gemini 3.8 Flash1871
GPT-5.4 Pro1850
GLM-5.3 Flash1843
Opus 4.81841
Gemini 3.7 Flash1828
GLM-5.31818
Muse Spark 1.31802
Kimi K31769
GPT-5.6 Luna Pro1762
GPT-6 Luna Pro1761
Gemini 3.1 Pro Preview1732
GPT-5.41721
Gemini 3.5 Flash1710
Gemini 3.6 Flash1685
Sonnet 51681
Opus 4.71664
GPT-5.3-Codex1634
Muse Spark 1.21624
GLM-5.21620
GPT-5.2 Pro1608
Grok 4.51605
GLM-5.11571
DeepSeek-V4-Flash-07311556
Opus 4.61541
GPT-5.21525
DeepSeek-V4-Pro1509
Kimi K2.61494
GPT-5.4 Mini1454
Gemini 3 Pro Preview1442
Loading Atlas data…