Atlas

Benchmarks

← All benchmarks

Snorkel WorkplaceAgents

Professional Work

Pass@1 on 200 professional work-product tasks spanning 96 occupations in 19 sectors. Harbor evaluates deliverables with deterministic checks and substantive rubric criteria; formatting cannot dominate the score. Native effort and agent scaffold are not disclosed.

Top models (higher is better)

ModelScore
Grok 4.617.1
Claude Fable 5.115.8
Grok 4.715.2
Claude Opus 514.9
Claude Opus 5.514.4
GPT-6 Astra14.4
Kimi K313.0
GLM-5.312.8
Gemini 3.8 Flash11.2
Loading Atlas data…