Atlas

Benchmarks

← All benchmarks

AutomationBench-AA

Agents · 2026-07-06

AutomationBench-AA is Artificial Analysis' independent evaluation of Zapier's AutomationBench for cross-application SaaS workflow agents. It reports the average share of task objectives completed with no guardrail violations across 657 simulated SaaS workflow tasks spanning finance, HR, marketing, operations, sales, and support.

Top models (higher is better)

ModelScore
Kimi K352.7
Grok 4.551.4
GPT-5.6 Sol51.2
Gemini 3.6 Flash51.1
Claude Fable 548.6
Opus 4.848.5
GPT-5.6 Terra45.6
Muse Spark 1.142.8
Gemini 3.5 Flash42.6
GPT-5.6 Luna42.2
GPT-5.542.1
Sonnet 539.2
Gemini 3.1 Pro Preview37.5
Gemini 3.5 Flash-Lite32.7
GLM-5.227.8
Qwen3.7-Max25.6
Sonnet 4.623.8
Kimi K2.7 Code22.5
Qwen3.7-Plus20.4
Kimi K2.619.6
Nex-N2-Pro19.4
DeepSeek-V4-Pro18.8
Qwen3.6 Plus (2026-04-02)18.6
MiMo-V2.5-Pro16.9
DeepSeek-V4-Flash16.6
MiniMax M315.9
Mistral Medium 3.513.7
Qwen3.5 397B A17B12.2
Gemini 3.1 Flash-Lite Preview10.6
KAT-Coder-Pro V19.6
Haiku 4.59.6
Inkling-Small9.4
Grok 4.38.1
Step 3.7 Flash7.0
Nemotron 3 Ultra 550B A55B5.7
Ring 2.6 1T2.1
GPT-5.5 Instant (2026-06-25 hosted snapshot)1.4
Loading Atlas data…