Atlas

Benchmarks

← All benchmarks

SWE Atlas - Codebase QnA (pass@3)

Code · 2026-03-04

This aggregation of the SWE Atlas Codebase QnA track reports pass@3 on repository-level questions requiring runtime analysis and multi-file reasoning. It is kept separate from single-run resolve-rate observations on the parent track.

Top models (higher is better)

ModelScore
Opus 4.857.3
Macaron-V1-Venti49.5
GLM-5.248.9
GPT-5.545.4
MiniMax M337.9
Qwen3.7-Max22.6
Gemini 3.1 Pro Preview13.5
Loading Atlas data…