Ramp SWE-Bench
Code · 2026-06-11
Ramp SWE-Bench is a private, production-grounded coding benchmark built from real Ramp backend engineering tasks. It evaluates background coding agents using mini-swe-agent on 80 human-accepted tasks, with pass rate reported as the percentage of tasks resolved. The dashboard's task manifest lists all 80 tasks, of which two (tasks 35 and 75) are currently unpublished, so published runs report over a 78-task denominator; evaluation_item_count records the benchmark's canonical 80-task size.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 88.5 |
| Claude Opus 5 | 88.5 |
| Kimi K3 | 87.2 |
| Opus 4.7 | 83.3 |
| GPT-5.5 | 83.3 |
| GPT-5.6 Sol | 83.3 |
| GLM-5.2 | 82.1 |
| Grok 4.5 | 82.1 |
| Opus 4.6 | 80.8 |
| Kimi K2.7 Code | 80.8 |
| DeepSeek-V4-Flash-0731 | 79.5 |
| Opus 4.8 | 78.2 |
| GPT-5.6 Terra | 76.9 |
| Sonnet 5 | 75.6 |
| Gemini 3.1 Pro Preview | 74.4 |
| GPT-5.4 | 74.4 |
| GPT-5.6 Luna | 74.4 |
| Sonnet 4.6 | 73.1 |
| Kimi K2.6 | 73.1 |
| GLM-5.1 | 70.5 |
| Qwen3.6 Plus (2026-04-02) | 66.7 |
| DeepSeek-V4-Pro | 65.4 |
| Qwen3.7-Plus | 62.8 |
| GPT-5.4 Mini | 60.3 |
| Haiku 4.5 | 50.0 |
| GPT-5.4 Nano | 50.0 |
| GPT-4.1 | 15.4 |