AutomationBench (Public Set)
Agents · 2026-04-20
Pass rate on the 600-task public split of Zapier's AutomationBench, the task set the project ships publicly for research and experimentation: 100 tasks in each of sales, marketing, operations, support, finance and HR, with the repository's 200 additional 'simple' foundational tasks excluded from the score. Scoring is strict per-task pass/fail, credited only when every assertion passes. Distinct from AutomationBench (Overall), which the official Zapier leaderboard scores on a separate held-out private task set that follows the same domain distribution but is never released; public-split scores run materially higher and are not comparable with it.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 30.8 |
| GPT-5.6 Sol | 29.7 |
| Claude Fable 5 | 29.1 |
| Opus 4.8 | 27.2 |
| GPT-5.5 | 22.7 |
| GLM-5.2 | 12.9 |