Atlas

Benchmarks

← All benchmarks

Snorkel WorkplaceAgents pass@5

Professional Work

Published pass@5 on the 200-task WorkplaceAgents release; a correlated diagnostic of the pass@1 leaderboard.

Top models (higher is better)

ModelScore
Claude Opus 534.7
Claude Fable 5.130.8
Grok 4.729.4
Grok 4.628.6
Kimi K328.2
Claude Opus 5.526.4
Gemini 3.8 Flash23.7
GPT-6 Astra23.7
GLM-5.323.6
Loading Atlas data…