Atlas

Benchmarks

← All benchmarks

ReactBench (39 tasks): Writing React

Code

Writing React subset of the revised ReactBench panel; 23 feature or fix tasks with held-out behavioral checks and React Doctor verification, five trials per task.

Top models (higher is better)

ModelScore
Claude Fable 561.7
Claude Fable 5.160.9
Claude Opus 555.7
GPT-5.6 Sol50.4
GPT-5.6 Terra50.4
GLM-5.3 Flash44.3
GPT-5.6 Luna42.6
Opus 4.840.9
Kimi K340.9
Grok 4.639.1
Sonnet 537.4
Gemini 3.8 Flash36.5
GLM-5.336.5
DeepSeek V4 Pro 081334.8
Grok 4.533.0
Gemini 3.1 Pro Preview30.4
GLM-5.230.4
Gemini 3.5 Flash27.8
Kimi K2.7 Code27.8
DeepSeek-V4-Flash-073127.0
Muse Spark 1.126.1
Muse Spark 1.225.2
Composer 2.519.1
Inkling-Small9.6
Loading Atlas data…