Atlas

Benchmarks

← All benchmarks

AutomationBench-AA 1.0.6

Agents

Artificial Analysis' AutomationBench 1.0.6 series, identified by automation_bench_1_0_6 in its canonical data schema. The 657-task implementation reports partial objective completion with a zero score for guardrail violations, distinct from earlier source revisions.

Top models (higher is better)

ModelScore
Claude Opus 5.569.5
DeepSeek V4.1 Flash68.9
GPT-6 Astra68.5
Grok 4.667.0
Grok 4.765.6
GLM-5.362.2
Gemini 3.7 Flash62.0
GPT-6 Sol61.7
Gemini 3.8 Flash60.9
GLM-5.3 Flash60.4
GPT-5.6 Sol60.1
GPT-5.6 Terra59.6
Claude Fable 5.159.4
MiMo-V2.6-Pro58.6
Kimi K358.3
Grok 4.557.9
Muse Spark 1.357.9
Qwen3.8 2.4T A95B57.2
DeepSeek V4 Pro 081356.7
Claude Opus 556.6
DeepSeek-V4-Pro56.3
Qwen3.8 Max (0902)56.2
Qwen3.8-Flash-Next55.9
Claude Fable 554.1
DeepSeek-V4-Flash-073154.0
GPT-6 Luna53.2
Gemini 3.6 Flash53.0
Step 5 Preview51.0
Agnes 3.0 Flash (hosted)50.7
GPT-5.6 Luna50.2
Qwen3.8-Max49.2
Qwen3.8 27B48.2
DeepSeek V4 Flash Vision Exp47.5
GPT-5.547.3
Opus 4.845.6
Gemini 3.5 Flash42.1
Muse Spark 1.240.6
Muse Spark 1.138.8
Agnes 2.5 Pro Alpha38.4
K2 Horizon 375B A23B37.2
Loading Atlas data…