Kimi Code Bench v2
Code · 2026-06-11
In-house coding-agent benchmark for production software tasks.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 76.9 |
| Kimi K3 | 72.9 |
| Opus 4.8 | 71.7 |
| GPT-5.5 | 69.0 |
| GPT-5.6 Sol | 64.8 |
| GLM-5.2 | 64.2 |
Code · 2026-06-11
In-house coding-agent benchmark for production software tasks.
| Model | Score |
|---|---|
| Claude Fable 5 | 76.9 |
| Kimi K3 | 72.9 |
| Opus 4.8 | 71.7 |
| GPT-5.5 | 69.0 |
| GPT-5.6 Sol | 64.8 |
| GLM-5.2 | 64.2 |