APEX-Accounting
Professional Work · 2026-07-31
Mercor and Ramp's 160 professional accounting tasks across 10 month-end-close worlds, scored against 2,186 expert-authored criteria. Mean Score averages the fraction of criteria passed per task using the Loop agent.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 61.8 |
| Claude Fable 5.1 | 61.0 |
| GPT-6 Astra | 57.9 |
| Claude Fable 5 | 56.4 |
| Claude Opus 5 | 54.5 |
| Muse Spark 1.3 | 53.5 |
| GLM-5.3 | 53.1 |
| Muse Spark 1.1 | 52.6 |
| Gemini 3.8 Flash | 51.7 |
| GPT-5.6 Sol | 51.5 |
| GPT-5.6 Sol Pro | 51.5 |
| Muse Spark 1.2 | 51.4 |
| Kimi K3 | 49.9 |
| GPT-5.5 | 48.2 |
| DeepSeek V4.1 Flash | 48.1 |
| Opus 4.8 | 48.0 |
| GPT-5.4 | 46.8 |
| Grok 4.7 | 46.7 |
| Grok 4.5 | 43.7 |
| Opus 4.6 | 43.3 |
| GLM-5.2 | 42.7 |
| GPT-5.6 Terra | 40.3 |
| Kimi K2.7 Code | 39.6 |
| MiniMax M3 | 39.1 |
| GPT-5.6 Luna | 38.0 |
| Gemini 3.1 Pro Preview | 34.0 |
| Inkling | 24.4 |
| Qwen3.5 397B A17B | 24.4 |
| Nemotron 3 Ultra 550B A55B (NVFP4) | 22.9 |
| DeepSeek-V3.2 | 20.7 |
| GLM-5.1 | 18.6 |
| GLM-5.3 Flash | 10.8 |
| gpt-oss-120b | 2.1 |