Atlas

Benchmarks

← All benchmarks

Vibe Code Bench 1–100

Code · 2026-09-22

Fifty held-out web-development scenarios with up to ten dependent feature requests and a shared ten-hour budget per scenario. Modified OpenHands v1 agent uses Supabase, Stripe and MailHog services. Score is the mean fraction of planned consecutive iterations completed perfectly before the first failure; it is not binary task accuracy. Fifty validation scenarios are excluded.

Top models (higher is better)

ModelScore
Claude Opus 5.530.4
Claude Opus 528.5
Claude Fable 5.128.0
GPT-6 Astra27.6
GPT-5.6 Luna22.6
Muse Spark 1.320.5
GPT-5.6 Sol20.0
GLM-5.320.0
Gemini 3.8 Flash18.8
Kimi K318.2
DeepSeek V4 Pro 081317.5
DeepSeek V4.1 Flash16.4
GLM-5.3 Flash16.0
GPT-5.6 Terra14.8
Grok 4.614.8
Sonnet 513.8
MiniMax M39.2
Inkling7.3
Gemini 3.1 Pro Preview6.7
Loading Atlas data…