Atlas

Benchmarks

← All benchmarks

FrontierFinance v3.1

Professional Work · 2026-07-06

Samaya AI's FrontierFinance v3.1 evaluates finance agents on 220 open-ended, dated research queries across six investor workflows, using 11,543 expert-authored rubric criteria. The headline metric is rubric qualification rate macro-averaged across queries after three-model majority-vote judging.

Top models (higher is better)

ModelScore
Claude Fable 549.2
GPT-5.6 Sol46.8
Opus 4.845.0
GPT-5.543.5
GLM-5.242.8
DeepSeek-V4-Pro40.5
Kimi K2.632.3
Gemini 3.1 Pro Preview30.7
Loading Atlas data…