Atlas

Benchmarks

← All benchmarks

SWE-bench Verified Epoch-Compatible 484-Task Subset

Code · 2026-02-01

The 484-task subset of SWE-bench Verified that passes on Epoch AI's evaluation infrastructure. Sixteen tasks from the 500-task parent benchmark are excluded, and individual runs may have fewer completions, so these scores are not directly comparable to full-set results.

Top models (higher is better)

ModelScore
Opus 4.783.5
GPT-5.580.6
Gemini 3.5 Flash79.3
GLM-5.278.7
DeepSeek-V4-Pro77.6
Qwen3.7-Max77.3
GPT-5.476.9
Opus 4.576.7
Kimi K2.676.7
Qwen3.6-Max-Preview76.7
Gemini 3.1 Pro Preview75.6
Gemini 3 Flash Preview75.4
Sonnet 4.675.2
GPT-5.3-Codex74.8
GLM-5.174.2
GPT-5.273.8
Kimi K2.573.8
GPT-573.6
Opus 4.173.3
Gemini 3 Pro Preview72.9
GLM-572.1
Sonnet 4.571.3
Opus 470.7
GPT-5 Mini64.7
o362.3
Claude 3.7 Sonnet61.0
Qwen3.6 Plus (2026-04-02)57.9
Gemini 2.5 Pro57.6
GPT-4.148.5
GPT-4o (2024-11-20)31.0
Loading Atlas data…