AA-AnalystAgent (pass@1)
Professional Work
AA-AnalystAgent accuracy averaged over five attempts for each of 80 private quantitative analysis questions. Uses the same Stirrup harness as the pass^5 headline.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.7 Flash | 70.5 |
| GPT-5.5 | 66.3 |
| Claude Fable 5.1 | 65.8 |
| Gemini 3.1 Pro Preview | 64.0 |
| Claude Opus 5 | 63.8 |
| Opus 4.8 | 63.5 |
| Sonnet 5 | 61.5 |
| GPT-5.6 Sol | 61.3 |
| Grok 4.6 | 60.8 |
| GPT-6 Astra | 60.0 |
| Claude Fable 5 | 59.8 |
| Opus 4.7 | 59.5 |
| Gemini 3.5 Flash | 58.5 |
| Kimi K3 | 57.8 |
| Inkling-Small | 57.0 |
| Grok 4.5 | 54.0 |
| Inkling | 52.3 |
| Qwen3.7-Max | 48.3 |
| DeepSeek-V4-Flash | 47.3 |
| MiMo-V2.5-Pro | 47.0 |
| Sonnet 4.6 | 46.8 |
| DeepSeek-V4-Pro | 45.0 |
| MiniMax M3 | 44.0 |
| GPT-5.4 Mini | 41.3 |
| Ling 3.0 Flash Fin | 40.8 |
| Grok 4.3 | 31.8 |
| Haiku 4.5 | 30.0 |
| Mistral Medium 3.5 | 30.0 |
| MiniMax M2.7 | 27.0 |
| Nemotron 3 Ultra 550B A55B | 25.0 |
| Gemini 3.1 Flash-Lite Preview | 24.0 |
| Mistral Small 4 | 7.3 |