Atlas

Benchmarks

← All benchmarks

ReactBench v1 Fix React

Code · 2026-07-14

The 24-task Fix React subset of ReactBench v1 asks coding agents to remove targeted React problems without being shown the React Doctor findings. A rollout passes only when the targeted issues are removed, no new graded issue is introduced, and existing behavior remains intact.

Top models (higher is better)

ModelScore
GPT-5.6 Sol50.8
GPT-5.6 Terra40.8
Grok 4.535.8
Claude Opus 532.5
GPT-5.6 Luna32.5
Claude Fable 526.7
GLM-5.224.2
Muse Spark 1.121.7
Opus 4.820.8
Kimi K319.2
Sonnet 518.3
Gemini 3.5 Flash18.3
Gemini 3.1 Pro Preview16.7
Composer 2.514.2
Kimi K2.7 Code14.2
Inkling7.5
Loading Atlas data…