Atlas

Benchmarks

← All benchmarks

ReactBench v1

Code · 2026-07-14

ReactBench v1 evaluates coding agents on 51 realistic React tasks from 47 open-source repositories. It combines 27 Write React and 24 Fix React tasks, grading each rollout with hidden behavioral tests and a pinned React Doctor gate; published scores average pass@1 over five trials per task.

Top models (higher is better)

ModelScore
GPT-5.6 Terra53.3
GPT-5.6 Sol52.5
Claude Opus 549.0
Claude Fable 547.5
GPT-5.6 Luna43.9
Grok 4.540.4
Opus 4.836.5
Kimi K332.9
GLM-5.232.5
Sonnet 530.6
Muse Spark 1.129.0
Gemini 3.5 Flash26.3
Gemini 3.1 Pro Preview23.5
Kimi K2.7 Code23.5
Composer 2.521.2
Inkling11.0
Loading Atlas data…