Atlas

Benchmarks

← All benchmarks

Snorkel Digital Work Index

Indexes

Weighted arithmetic average of eight frontier benchmark pass rates. Weights are normalized geometric means of wage coverage and evaluated task volume, with industry adjustments and shared-occupation allocation. The release maps to 29 occupations. Diagnostic only because component benchmarks are imported independently; native invocation controls and model-specific standard errors are not disclosed.

Top models (higher is better)

ModelScore
Claude Fable 5.123.6
Claude Opus 5.523.6
GPT-6 Astra23.0
Claude Opus 521.8
Grok 4.620.5
Grok 4.719.3
Gemini 3.8 Flash16.9
Kimi K316.8
GLM-5.316.4
Muse Spark 1.316.3
Loading Atlas data…