APEX (Mercor)
Professional Work · 2025-10-01
Mercor's APEX evaluates models on realistic professional tasks authored by veteran practitioners across four occupations -- investment banking associate, management consultant, big law associate, and primary care physician -- with 100 tasks each. Every task is answered eight times and graded by a judge model, and the reported score is the mean.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.6 | 14.7 |
| GPT-5.4 | 14.5 |
| GPT-5.2-Codex | 13.6 |
| GPT-5.3-Codex | 13.5 |
| Gemini 3.1 Pro Preview | 13.3 |
| GPT-5 | 12.9 |
| GPT-5.2 | 11.9 |
| GPT-5-Codex | 11.7 |
| GPT-5.1-Codex | 11.5 |
| Sonnet 4.6 | 10.5 |
| Gemini 3 Flash Preview | 8.9 |
| Gemini 3 Pro Preview | 8.8 |
| o3 | 8.6 |
| Grok 4 | 8.2 |
| Opus 4.5 | 7.7 |
| Grok 4.1 Fast | 7.7 |
| GLM-5 | 7.3 |
| GLM-4.7 | 6.2 |