Atlas

Benchmarks

← All benchmarks

SWE-bench Multimodal

Code

SWE-bench Multimodal extends the SWE-bench issue-resolution format with visual context such as screenshots and design mockups embedded in the issue description, testing whether coding agents generalize to visual software domains.

Top models (higher is better)

ModelScore
Claude Opus 559.4
Claude Mythos Preview59.0
Claude Mythos 554.9
Claude Fable 554.1
Opus 4.838.4
Opus 4.734.5
Sonnet 528.1
Opus 4.627.1
Loading Atlas data…