Atlas

Benchmarks

← All benchmarks

ReactBench (39 tasks): Fixing React

Code

Fixing React subset of the revised ReactBench panel; 16 refactoring tasks requiring every target React Doctor finding to be removed without regressions, five trials per task.

Top models (higher is better)

ModelScore
GPT-5.6 Sol41.3
GPT-5.6 Terra26.3
Grok 4.526.3
Grok 4.625.0
Claude Fable 5.123.8
Claude Opus 523.8
Claude Fable 522.5
Gemini 3.8 Flash22.5
DeepSeek V4 Pro 081322.5
GPT-5.6 Luna21.3
Sonnet 521.3
Muse Spark 1.120.0
Gemini 3.1 Pro Preview18.8
Opus 4.817.5
Gemini 3.5 Flash17.5
Kimi K316.3
GLM-5.315.0
GLM-5.215.0
GLM-5.3 Flash12.5
Muse Spark 1.211.3
Kimi K2.7 Code10.0
DeepSeek-V4-Flash-07316.3
Composer 2.55.0
Inkling-Small2.5
Loading Atlas data…