ReactBench (39 tasks)
Code
Revised 39-task ReactBench panel: 23 writing and 16 fixing tasks. Five trials per task, passing both behavioral tests and the React Doctor no-new-issues gate. Kept separate from the earlier 51-task panel.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 46.7 |
| Claude Fable 5.1 | 45.6 |
| Claude Fable 5 | 43.6 |
| Claude Opus 5 | 42.1 |
| GPT-5.6 Terra | 40.5 |
| GPT-5.6 Luna | 33.8 |
| Grok 4.6 | 32.3 |
| GLM-5.3 Flash | 31.3 |
| Gemini 3.8 Flash | 30.8 |
| Opus 4.8 | 30.8 |
| Kimi K3 | 30.8 |
| DeepSeek V4 Pro 0813 | 29.7 |
| Sonnet 5 | 29.7 |
| Grok 4.5 | 28.7 |
| GLM-5.3 | 27.7 |
| Gemini 3.1 Pro Preview | 25.6 |
| GLM-5.2 | 24.1 |
| Gemini 3.5 Flash | 23.6 |
| Muse Spark 1.1 | 23.1 |
| Kimi K2.7 Code | 20.5 |
| Muse Spark 1.2 | 19.5 |
| DeepSeek-V4-Flash-0731 | 18.5 |
| Composer 2.5 | 13.3 |
| Inkling-Small | 6.7 |