Atlas

Benchmarks

← All benchmarks

Vibe Code Bench v1.1 (Vals Index 22-task subset)

Code · 2026-07-01

Vibe Code Bench v1.1 evaluated on the 22 representative app-building tasks selected for the Vals Index across UI, data, and workflow patterns. It remains a protocol-specific subset of the full OpenHands-based benchmark.

Top models (higher is better)

ModelScore
Claude Fable 593.5
Kimi K391.3
Opus 4.889.9
GPT-5.6 Sol89.3
Sonnet 586.9
Claude Opus 585.7
GPT-5.6 Luna81.1
Opus 4.777.7
GPT-5.577.1
GLM-5.272.9
GPT-5.6 Terra72.8
Grok 4.572.8
Muse Spark 1.172.4
Gemini 3.6 Flash57.9
MiniMax M356.5
Sonnet 4.656.2
DeepSeek-V4-Pro55.9
Gemini 3.5 Flash54.7
Qwen3.7-Max52.9
MiMo-V2.548.1
Qwen3.7-Plus46.9
GPT-5.4 Mini42.5
Kimi K2.642.3
MiMo-V2.5-Pro36.3
Gemini 3.1 Pro Preview34.2
GLM-5.132.8
Gemini 3.5 Flash-Lite32.8
GPT-5.4 Nano30.0
Qwen3.6 Plus (2026-04-02)28.3
Gemini 3 Flash Preview22.0
Inkling-Small19.0
Grok 4.315.5
Inkling13.3
MiniMax M2.711.7
Haiku 4.59.9
Nemotron 3 Ultra 550B A55B7.1
Laguna XS.25.5
Laguna M.15.3
Grok 4.201.5
Mistral Medium 3.51.2
Loading Atlas data…