Atlas

Benchmarks

← All benchmarks

ReactBench v1 Write React

Code · 2026-07-14

The 27-task Write React subset of ReactBench v1 asks coding agents to implement real features or fixes from open-source pull requests. A rollout passes only when held-out behavioral tests succeed and the changed surface introduces no new graded React Doctor issue.

Top models (higher is better)

ModelScore
Claude Fable 565.9
GPT-5.6 Terra64.4
Claude Opus 563.7
GPT-5.6 Sol57.8
GPT-5.6 Luna54.1
Opus 4.850.4
Grok 4.548.1
GLM-5.247.4
Kimi K345.2
Sonnet 542.2
Muse Spark 1.135.6
Gemini 3.5 Flash33.3
Gemini 3.1 Pro Preview32.6
Kimi K2.7 Code31.9
Composer 2.527.4
Inkling14.1
Loading Atlas data…