Atlas

Benchmarks

← All benchmarks

APEX (Mercor)

Professional Work · 2025-10-01

Mercor's APEX evaluates models on realistic professional tasks authored by veteran practitioners across four occupations -- investment banking associate, management consultant, big law associate, and primary care physician -- with 100 tasks each. Every task is answered eight times and graded by a judge model, and the reported score is the mean.

Top models (higher is better)

ModelScore
Opus 4.614.7
GPT-5.414.5
GPT-5.2-Codex13.6
GPT-5.3-Codex13.5
Gemini 3.1 Pro Preview13.3
GPT-512.9
GPT-5.211.9
GPT-5-Codex11.7
GPT-5.1-Codex11.5
Sonnet 4.610.5
Gemini 3 Flash Preview8.9
Gemini 3 Pro Preview8.8
o38.6
Grok 48.2
Opus 4.57.7
Grok 4.1 Fast7.7
GLM-57.3
GLM-4.76.2
Loading Atlas data…