Atlas

Benchmarks

← All benchmarks

APEX-Accounting

Professional Work · 2026-07-31

Mercor and Ramp's 160 professional accounting tasks across 10 month-end-close worlds, scored against 2,186 expert-authored criteria. Mean Score averages the fraction of criteria passed per task using the Loop agent.

Top models (higher is better)

ModelScore
Claude Opus 5.561.8
Claude Fable 5.161.0
GPT-6 Astra57.9
Claude Fable 556.4
Claude Opus 554.5
Muse Spark 1.353.5
GLM-5.353.1
Muse Spark 1.152.6
Gemini 3.8 Flash51.7
GPT-5.6 Sol51.5
GPT-5.6 Sol Pro51.5
Muse Spark 1.251.4
Kimi K349.9
GPT-5.548.2
DeepSeek V4.1 Flash48.1
Opus 4.848.0
GPT-5.446.8
Grok 4.746.7
Grok 4.543.7
Opus 4.643.3
GLM-5.242.7
GPT-5.6 Terra40.3
Kimi K2.7 Code39.6
MiniMax M339.1
GPT-5.6 Luna38.0
Gemini 3.1 Pro Preview34.0
Inkling24.4
Qwen3.5 397B A17B24.4
Nemotron 3 Ultra 550B A55B (NVFP4)22.9
DeepSeek-V3.220.7
GLM-5.118.6
GLM-5.3 Flash10.8
gpt-oss-120b2.1
Loading Atlas data…