Vibe Code Bench 1–100
Code · 2026-09-22
Fifty held-out web-development scenarios with up to ten dependent feature requests and a shared ten-hour budget per scenario. Modified OpenHands v1 agent uses Supabase, Stripe and MailHog services. Score is the mean fraction of planned consecutive iterations completed perfectly before the first failure; it is not binary task accuracy. Fifty validation scenarios are excluded.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 30.4 |
| Claude Opus 5 | 28.5 |
| Claude Fable 5.1 | 28.0 |
| GPT-6 Astra | 27.6 |
| GPT-5.6 Luna | 22.6 |
| Muse Spark 1.3 | 20.5 |
| GPT-5.6 Sol | 20.0 |
| GLM-5.3 | 20.0 |
| Gemini 3.8 Flash | 18.8 |
| Kimi K3 | 18.2 |
| DeepSeek V4 Pro 0813 | 17.5 |
| DeepSeek V4.1 Flash | 16.4 |
| GLM-5.3 Flash | 16.0 |
| GPT-5.6 Terra | 14.8 |
| Grok 4.6 | 14.8 |
| Sonnet 5 | 13.8 |
| MiniMax M3 | 9.2 |
| Inkling | 7.3 |
| Gemini 3.1 Pro Preview | 6.7 |