Atlas

Benchmarks

← All benchmarks

DAYJOB: Finance

Professional Work · 2026-09-23

Surge AI DAYJOB: Finance evaluates professional judgment on 80 expert-designed finance assignments in rich file environments. Each task requires complete rubric satisfaction. The leaderboard averages task rewards over five attempts, excludes errored trials, and reports the unweighted mean across tasks. Runs use the DAYJOB OpenHands SDK harness in Harbor with no browsing or package downloads and Claude Opus 4.8 agentic grading.

Top models (higher is better)

ModelScore
Claude Opus 5.523.9
GPT-6 Astra21.5
Claude Fable 5.119.8
Muse Spark 1.314.8
Grok 4.714.5
Grok 4.612.8
Claude Opus 511.3
GPT-6 Sol9.3
Claude Fable 59.0
GLM-5.36.0
GPT-5.6 Sol5.5
Kimi K33.8
Gemini 3.7 Flash2.8
Gemini 3.8 Flash2.8
Sonnet 52.3
GLM-5.3 Flash2.3
GPT-6 Luna2.3
Muse Spark 1.22.3
GPT-5.6 Luna1.8
GPT-5.6 Terra1.8
Hy30.3
Gemini 3.1 Pro Preview0.0
Inkling0.0
Kimi K2.7 Code0.0
Mistral Large 3 675B Instruct 25120.0
Muse Glimmer 30B0.0
Loading Atlas data…