SWE-bench Multilingual
Code · 2025-05-06
SWE-bench Multilingual evaluates software-engineering agents on 300 curated issue-resolution tasks drawn from 42 repositories across nine programming languages. Agents modify a repository snapshot to resolve a real GitHub issue, and success requires both issue-specific and regression tests to pass.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos 5 | 92.2 |
| Claude Opus 5 | 89.5 |
| Claude Mythos Preview | 87.3 |
| Claude Fable 5 | 86.6 |
| Opus 4.8 | 84.4 |
| Opus 4.7 | 80.5 |
| Laguna S 2.1 | 78.5 |
| Sonnet 5 | 78.3 |
| Qwen3.7-Max | 78.3 |
| Opus 4.6 | 77.8 |
| Opus 4.5 | 76.2 |
| DeepSeek-V4-Pro | 76.2 |
| Sonnet 4.6 | 75.9 |
| Hy3 | 75.8 |
| Gemini 3 Flash Preview | 72.7 |
| GLM-5 | 69.7 |
| Gemini 3 Pro Preview | 68.7 |
| MiniMax M2.5 | 68.3 |
| Nemotron 3 Ultra 550B A55B | 67.7 |
| Kimi K2.5 | 67.3 |
| Sonnet 4.5 | 67.0 |
| GPT-5.2 (2025-12-11) | 66.7 |
| GPT-5.2-Codex | 66.3 |
| Haiku 4.5 | 64.7 |
| DeepSeek-V3.2 | 59.0 |
| GPT-5 Mini (2025-08-07) | 39.7 |