Atlas

Benchmarks

← All benchmarks

ReactBench (39 tasks)

Code

Revised 39-task ReactBench panel: 23 writing and 16 fixing tasks. Five trials per task, passing both behavioral tests and the React Doctor no-new-issues gate. Kept separate from the earlier 51-task panel.

Top models (higher is better)

ModelScore
GPT-5.6 Sol46.7
Claude Fable 5.145.6
Claude Fable 543.6
Claude Opus 542.1
GPT-5.6 Terra40.5
GPT-5.6 Luna33.8
Grok 4.632.3
GLM-5.3 Flash31.3
Gemini 3.8 Flash30.8
Opus 4.830.8
Kimi K330.8
DeepSeek V4 Pro 081329.7
Sonnet 529.7
Grok 4.528.7
GLM-5.327.7
Gemini 3.1 Pro Preview25.6
GLM-5.224.1
Gemini 3.5 Flash23.6
Muse Spark 1.123.1
Kimi K2.7 Code20.5
Muse Spark 1.219.5
DeepSeek-V4-Flash-073118.5
Composer 2.513.3
Inkling-Small6.7
Loading Atlas data…