App-Bench
Agents · 2025-10-25
App-Bench evaluates hosted app builders and coding assistants on six full-stack web-application tasks covering healthcare, real estate, finance, legal services, and education. Each tool receives one prompt per task, gets three independent one-shot attempts under its default Pro settings, and is scored by the best run using manually verified binary functional requirements; the aggregate is completed rubric points divided by all possible points.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.5 | 67.5 |
| Gemini 3 Pro Preview | 50.3 |
| GPT-5.1-Codex-Max | 38.4 |
| Gemini 2.5 Pro | 0.0 |