Atlas

Benchmarks

← All benchmarks

Ramp SWE-Bench

Code · 2026-06-11

Ramp SWE-Bench is a private, production-grounded coding benchmark built from real Ramp backend engineering tasks. It evaluates background coding agents using mini-swe-agent on 80 human-accepted tasks, with pass rate reported as the percentage of tasks resolved. The dashboard's task manifest lists all 80 tasks, of which two (tasks 35 and 75) are currently unpublished, so published runs report over a 78-task denominator; evaluation_item_count records the benchmark's canonical 80-task size.

Top models (higher is better)

ModelScore
Claude Fable 588.5
Claude Opus 588.5
Kimi K387.2
Opus 4.783.3
GPT-5.583.3
GPT-5.6 Sol83.3
GLM-5.282.1
Grok 4.582.1
Opus 4.680.8
Kimi K2.7 Code80.8
DeepSeek-V4-Flash-073179.5
Opus 4.878.2
GPT-5.6 Terra76.9
Sonnet 575.6
Gemini 3.1 Pro Preview74.4
GPT-5.474.4
GPT-5.6 Luna74.4
Sonnet 4.673.1
Kimi K2.673.1
GLM-5.170.5
Qwen3.6 Plus (2026-04-02)66.7
DeepSeek-V4-Pro65.4
Qwen3.7-Plus62.8
GPT-5.4 Mini60.3
Haiku 4.550.0
GPT-5.4 Nano50.0
GPT-4.115.4
Loading Atlas data…