Atlas

Benchmarks

← All benchmarks

JobBench Main

Professional Work · 2026-04-17

JobBench Main is the 65-task full-difficulty split spanning 35 occupations, 569 weighted rubric chains, and 2,066 binary criteria. Its official score is the unweighted macro-average of each task's normalized passed-rubric weight, with no partial credit within a chain.

Top models (higher is better)

ModelScore
Claude Fable 557.4
Muse Spark 1.154.7
Kimi K352.9
Opus 4.848.4
GPT-5.6 Sol46.5
Opus 4.744.5
GLM-5.243.4
GPT-5.538.3
Sonnet 4.636.6
GPT-5.432.2
Gemini 3.5 Flash31.5
GPT-5.226.6
Sonnet 4.520.7
Gemini 3.1 Pro Preview15.9
Loading Atlas data…