Atlas

Benchmarks

← All benchmarks

AA-AnalystAgent (pass@1)

Professional Work

AA-AnalystAgent accuracy averaged over five attempts for each of 80 private quantitative analysis questions. Uses the same Stirrup harness as the pass^5 headline.

Top models (higher is better)

ModelScore
Gemini 3.7 Flash70.5
GPT-5.566.3
Claude Fable 5.165.8
Gemini 3.1 Pro Preview64.0
Claude Opus 563.8
Opus 4.863.5
Sonnet 561.5
GPT-5.6 Sol61.3
Grok 4.660.8
GPT-6 Astra60.0
Claude Fable 559.8
Opus 4.759.5
Gemini 3.5 Flash58.5
Kimi K357.8
Inkling-Small57.0
Grok 4.554.0
Inkling52.3
Qwen3.7-Max48.3
DeepSeek-V4-Flash47.3
MiMo-V2.5-Pro47.0
Sonnet 4.646.8
DeepSeek-V4-Pro45.0
MiniMax M344.0
GPT-5.4 Mini41.3
Ling 3.0 Flash Fin40.8
Grok 4.331.8
Haiku 4.530.0
Mistral Medium 3.530.0
MiniMax M2.727.0
Nemotron 3 Ultra 550B A55B25.0
Gemini 3.1 Flash-Lite Preview24.0
Mistral Small 47.3
Loading Atlas data…