Atlas

Benchmarks

← All benchmarks

Voratiq Coding Agent Leaderboard

Code · 2026-02-03

Voratiq Coding Agent Leaderboard ranks coding agents on real engineering work in production codebases. Agents compete head-to-head, and scores are Bradley-Terry ratings fitted to winner/loser pairs from whose code gets merged.

Top models (higher is better)

ModelScore
Claude Fable 52253
Opus 4.82030
GPT-5.51914
GLM-5.21889
GPT-5.41872
Opus 4.71740
Opus 4.61563
GPT-5.4 Mini1509
Opus 4.51439
Sonnet 4.61423
DeepSeek-V4-Pro1374
Kimi K2.7 Code1374
GPT-5.3-Codex-Spark1360
Sonnet 4.51262
DeepSeek-V4-Flash1229
Gemini 3.1 Pro Preview1226
Haiku 4.51220
Gemini 3 Flash Preview1204
Loading Atlas data…