Atlas

Benchmarks

← All benchmarks

Next.js AI Agent Evaluations

Agents · 2025-10-21

Next.js AI Agent Evaluations measures whether coding agents can complete 24 framework-specific generation, migration, caching, routing, rendering, and best-practice tasks. Its official leaderboard reports the task success rate for paired runs with and without bundled Next.js documentation supplied through AGENTS.md.

Top models (higher is better)

ModelScore
Kimi K396.0
Claude Fable 596.0
Composer 2.596.0
GLM-5.296.0
Opus 4.896.0
Grok 4.596.0
Kimi K2.7 Code96.0
MiniMax M396.0
GLM-5.196.0
Opus 4.796.0
Opus 4.696.0
Kimi K2.696.0
Sonnet 4.696.0
GPT-5.6 Sol92.0
GPT-5.3-Codex92.0
Gemini 3.1 Pro Preview92.0
Composer 292.0
GPT-5.488.0
GPT-5.5 Pro83.0
Gemini 3 Pro Preview83.0
Sonnet 4.583.0
GPT-5.2-Codex79.0
MiniMax M2.763.0
Kimi K2.558.0
Loading Atlas data…