Next.js AI Agent Evaluations
Agents · 2025-10-21
Next.js AI Agent Evaluations measures whether coding agents can complete 24 framework-specific generation, migration, caching, routing, rendering, and best-practice tasks. Its official leaderboard reports the task success rate for paired runs with and without bundled Next.js documentation supplied through AGENTS.md.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 96.0 |
| Claude Fable 5 | 96.0 |
| Composer 2.5 | 96.0 |
| GLM-5.2 | 96.0 |
| Opus 4.8 | 96.0 |
| Grok 4.5 | 96.0 |
| Kimi K2.7 Code | 96.0 |
| MiniMax M3 | 96.0 |
| GLM-5.1 | 96.0 |
| Opus 4.7 | 96.0 |
| Opus 4.6 | 96.0 |
| Kimi K2.6 | 96.0 |
| Sonnet 4.6 | 96.0 |
| GPT-5.6 Sol | 92.0 |
| GPT-5.3-Codex | 92.0 |
| Gemini 3.1 Pro Preview | 92.0 |
| Composer 2 | 92.0 |
| GPT-5.4 | 88.0 |
| GPT-5.5 Pro | 83.0 |
| Gemini 3 Pro Preview | 83.0 |
| Sonnet 4.5 | 83.0 |
| GPT-5.2-Codex | 79.0 |
| MiniMax M2.7 | 63.0 |
| Kimi K2.5 | 58.0 |