Snorkel Digital Work Index
Indexes
Weighted arithmetic average of eight frontier benchmark pass rates. Weights are normalized geometric means of wage coverage and evaluated task volume, with industry adjustments and shared-occupation allocation. The release maps to 29 occupations. Diagnostic only because component benchmarks are imported independently; native invocation controls and model-specific standard errors are not disclosed.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5.1 | 23.6 |
| Claude Opus 5.5 | 23.6 |
| GPT-6 Astra | 23.0 |
| Claude Opus 5 | 21.8 |
| Grok 4.6 | 20.5 |
| Grok 4.7 | 19.3 |
| Gemini 3.8 Flash | 16.9 |
| Kimi K3 | 16.8 |
| GLM-5.3 | 16.4 |
| Muse Spark 1.3 | 16.3 |