Atlas

Benchmarks

← All benchmarks

AA-AnalystAgent (pass^5)

Professional Work

Artificial Analysis' private 80-question spreadsheet and document analysis benchmark across 14 domains. The headline pass^5 is the fraction of questions answered correctly on all five independent attempts with the Stirrup agent harness.

Top models (higher is better)

ModelScore
Gemini 3.7 Flash60.0
Claude Fable 5.157.5
Claude Opus 553.8
GPT-6 Astra51.3
GPT-5.550.0
Claude Fable 548.8
GPT-5.6 Sol47.5
Sonnet 546.3
Opus 4.845.0
Gemini 3.5 Flash45.0
Opus 4.743.8
Gemini 3.1 Pro Preview41.3
Grok 4.641.3
Kimi K338.8
Grok 4.535.0
Inkling-Small27.5
DeepSeek-V4-Flash25.0
Inkling23.8
Sonnet 4.620.0
MiMo-V2.5-Pro20.0
DeepSeek-V4-Pro18.8
Qwen3.7-Max18.8
Ling 3.0 Flash Fin16.3
Haiku 4.515.0
GPT-5.4 Mini15.0
Mistral Medium 3.512.5
MiniMax M2.711.3
MiniMax M310.0
Gemini 3.1 Flash-Lite Preview8.8
Grok 4.38.8
Nemotron 3 Ultra 550B A55B6.3
Mistral Small 41.3
Loading Atlas data…