AutomationBench-AA
Agents · 2026-07-06
AutomationBench-AA is Artificial Analysis' independent evaluation of Zapier's AutomationBench for cross-application SaaS workflow agents. It reports the average share of task objectives completed with no guardrail violations across 657 simulated SaaS workflow tasks spanning finance, HR, marketing, operations, sales, and support.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 52.7 |
| Grok 4.5 | 51.4 |
| GPT-5.6 Sol | 51.2 |
| Gemini 3.6 Flash | 51.1 |
| Claude Fable 5 | 48.6 |
| Opus 4.8 | 48.5 |
| GPT-5.6 Terra | 45.6 |
| Muse Spark 1.1 | 42.8 |
| Gemini 3.5 Flash | 42.6 |
| GPT-5.6 Luna | 42.2 |
| GPT-5.5 | 42.1 |
| Sonnet 5 | 39.2 |
| Gemini 3.1 Pro Preview | 37.5 |
| Gemini 3.5 Flash-Lite | 32.7 |
| GLM-5.2 | 27.8 |
| Qwen3.7-Max | 25.6 |
| Sonnet 4.6 | 23.8 |
| Kimi K2.7 Code | 22.5 |
| Qwen3.7-Plus | 20.4 |
| Kimi K2.6 | 19.6 |
| Nex-N2-Pro | 19.4 |
| DeepSeek-V4-Pro | 18.8 |
| Qwen3.6 Plus (2026-04-02) | 18.6 |
| MiMo-V2.5-Pro | 16.9 |
| DeepSeek-V4-Flash | 16.6 |
| MiniMax M3 | 15.9 |
| Mistral Medium 3.5 | 13.7 |
| Qwen3.5 397B A17B | 12.2 |
| Gemini 3.1 Flash-Lite Preview | 10.6 |
| KAT-Coder-Pro V1 | 9.6 |
| Haiku 4.5 | 9.6 |
| Inkling-Small | 9.4 |
| Grok 4.3 | 8.1 |
| Step 3.7 Flash | 7.0 |
| Nemotron 3 Ultra 550B A55B | 5.7 |
| Ring 2.6 1T | 2.1 |
| GPT-5.5 Instant (2026-06-25 hosted snapshot) | 1.4 |