Atlas

Benchmarks

← All benchmarks

AA-Briefcase Elo

Agents · 2026-06-18

AA-Briefcase is an Artificial Analysis benchmark for long-horizon agentic knowledge-work projects. It evaluates agents on realistic multi-week business workflows with many linked tasks and thousands of input files, combining rubric and pairwise grading into an Elo-style score.

Top models (higher is better)

ModelScore
Claude Opus 51470
Sonnet 51385
Grok 4.51317
Opus 4.71280
MiniMax M31108
GPT-5.51099
Sonnet 4.61076
GLM-5.1971
Gemini 3.6 Flash962
DeepSeek-V4-Pro931
Qwen3.7-Max914
MiMo-V2.5-Pro879
Nemotron 3 Ultra 550B A55B874
Gemini 3.5 Flash873
GPT-5.3-Codex869
Muse Spark 1.1868
DeepSeek-V4-Flash833
Kimi K2.6818
Qwen3.6 27B810
Grok 4.3759
GPT-5.4 Mini717
Muse Spark641
Gemini 3.5 Flash-Lite635
Haiku 4.5611
KAT-Coder-Pro V1598
Qwen3.5 397B A17B553
Mistral Medium 3.5516
Gemini 3.1 Pro Preview457
Gemma 4 31B IT373
Command A+ (May 2026, BF16)369
North Mini Code238
Gemini 3.1 Flash-Lite Preview230
Solar Pro 3138
K2 Think V258.8
gpt-oss-120b7.0
gpt-oss-20b-36.8
Nemotron 3 Super 120B A12B-95.3
Llama 4 Maverick Instruct-107
Loading Atlas data…