Atlas

Benchmarks

← All benchmarks

APEX-Agents-AA (avg@3)

Professional Work · 2026-01-20

Artificial Analysis' APEX-Agents-AA protocol evaluates 452 long-horizon professional-services tasks with three independent attempts per task (1,356 scored attempts) using the Stirrup/Archipelago agent harness. The reported pass@1 is the repeated-run mean.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash47.1
GPT-5.6 Terra38.9
GPT-5.537.7
GPT-5.6 Luna35.8
GLM-5.233.7
GPT-5.433.3
Opus 4.633.0
Gemini 3.1 Pro Preview32.0
Kimi K2.628.5
GPT-5.4 Mini28.2
Sonnet 4.628.0
Gemini 3 Flash Preview27.7
GPT-5.4 Nano24.9
DeepSeek-V4-Pro24.3
Qwen3.7-Plus22.4
Grok 4.317.0
Qwen3.5 397B A17B15.3
Step 3.7 Flash14.8
DeepSeek-V3.214.5
GLM-514.5
Grok 4.2014.2
Gemini 3.1 Flash-Lite Preview12.2
Kimi K2.511.5
MiniMax M2.710.6
gpt-oss-120b3.1
MiMo-V2.5-Pro2.4
Nemotron 3 Super 120B A12B1.8
gpt-oss-20b0.7
Loading Atlas data…