CUA-bench
Agents · 2026-09-18
Mean ordered-milestone progress across six computer-use game tasks. Each task has a three-hour budget and a 100-point milestone ladder; one campaign per model. This is progress, not binary task accuracy.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 19.2 |
| Claude Fable 5.1 | 13.2 |
| Claude Opus 5 | 9.0 |
| GPT-5.6 Sol | 8.3 |
| Gemini 3.8 Flash | 4.2 |