Atlas

Benchmarks

← All benchmarks

SnorkelRevOps

Professional Work

Pass@1 on 200 simulated B2B SaaS revenue-operations tasks across CRM, CPQ, billing, compensation, marketing, usage, finance and governance systems, including control-sensitive state changes. Native effort and agent scaffold are not disclosed.

Top models (higher is better)

ModelScore
Grok 4.615.8
Claude Fable 5.114.0
Kimi K313.1
Muse Spark 1.312.6
Claude Opus 5.512.5
Claude Opus 512.2
GLM-5.311.8
Grok 4.711.3
GPT-6 Astra10.9
Gemini 3.8 Flash5.4
Loading Atlas data…