Atlas

Benchmarks

← All benchmarks

SnorkelFinance 2.0 pass@5

Professional Work

Published pass@5 on the 200-task SnorkelFinance 2.0 workflow release; a correlated diagnostic of the pass@1 leaderboard.

Top models (higher is better)

ModelScore
GLM-5.342.3
Kimi K340.9
Grok 4.637.4
Grok 4.735.0
Claude Fable 5.134.2
Claude Opus 533.5
Claude Opus 5.532.0
Muse Spark 1.324.0
GPT-6 Astra19.3
Gemini 3.8 Flash16.0
Loading Atlas data…