ReactBench (39 tasks): Fixing React
Code
Fixing React subset of the revised ReactBench panel; 16 refactoring tasks requiring every target React Doctor finding to be removed without regressions, five trials per task.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 41.3 |
| GPT-5.6 Terra | 26.3 |
| Grok 4.5 | 26.3 |
| Grok 4.6 | 25.0 |
| Claude Fable 5.1 | 23.8 |
| Claude Opus 5 | 23.8 |
| Claude Fable 5 | 22.5 |
| Gemini 3.8 Flash | 22.5 |
| DeepSeek V4 Pro 0813 | 22.5 |
| GPT-5.6 Luna | 21.3 |
| Sonnet 5 | 21.3 |
| Muse Spark 1.1 | 20.0 |
| Gemini 3.1 Pro Preview | 18.8 |
| Opus 4.8 | 17.5 |
| Gemini 3.5 Flash | 17.5 |
| Kimi K3 | 16.3 |
| GLM-5.3 | 15.0 |
| GLM-5.2 | 15.0 |
| GLM-5.3 Flash | 12.5 |
| Muse Spark 1.2 | 11.3 |
| Kimi K2.7 Code | 10.0 |
| DeepSeek-V4-Flash-0731 | 6.3 |
| Composer 2.5 | 5.0 |
| Inkling-Small | 2.5 |