Snorkel WorkplaceAgents
Professional Work
Pass@1 on 200 professional work-product tasks spanning 96 occupations in 19 sectors. Harbor evaluates deliverables with deterministic checks and substantive rubric criteria; formatting cannot dominate the score. Native effort and agent scaffold are not disclosed.
Top models (higher is better)
| Model | Score |
|---|---|
| Grok 4.6 | 17.1 |
| Claude Fable 5.1 | 15.8 |
| Grok 4.7 | 15.2 |
| Claude Opus 5 | 14.9 |
| Claude Opus 5.5 | 14.4 |
| GPT-6 Astra | 14.4 |
| Kimi K3 | 13.0 |
| GLM-5.3 | 12.8 |
| Gemini 3.8 Flash | 11.2 |