Atlas

Benchmarks

← All benchmarks

AA-AnalystAgent (pass@5)

Professional Work

Fraction of AA-AnalystAgent's 80 quantitative analysis questions answered correctly at least once across five independent attempts.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview81.3
Inkling-Small80.0
Opus 4.878.8
GPT-5.578.8
Inkling78.8
Sonnet 577.5
Gemini 3.7 Flash77.5
Gemini 3.5 Flash76.3
Grok 4.576.3
Grok 4.676.3
DeepSeek-V4-Pro75.0
Qwen3.7-Max75.0
Claude Opus 573.8
MiniMax M373.8
Claude Fable 5.172.5
Opus 4.772.5
Kimi K372.5
Sonnet 4.671.3
DeepSeek-V4-Flash70.0
GPT-5.6 Sol70.0
Claude Fable 568.8
MiMo-V2.5-Pro68.8
GPT-6 Astra67.5
Ling 3.0 Flash Fin67.5
GPT-5.4 Mini62.5
Grok 4.357.5
Haiku 4.551.3
MiniMax M2.748.8
Mistral Medium 3.547.5
Gemini 3.1 Flash-Lite Preview43.8
Nemotron 3 Ultra 550B A55B43.8
Mistral Small 423.8
Loading Atlas data…