Atlas

Benchmarks

← All benchmarks

Ruby LLM Program Fixer — Parking Garage

Code · 2025-08-01

The Parking Garage Program Fixer task asks a model to repair a broken Ruby parking-garage implementation against 39 tests. Its composite is 90% test success plus 10% RuboCop quality, with each offense subtracting two quality points down to zero.

Top models (higher is better)

ModelScore
GPT-5.585.0
GPT-5.6 Sol80.1
GPT-5.6 Sol Pro79.9
GPT-5.477.0
GPT-5.6 Terra Pro77.0
Claude Opus 572.0
GPT-5.6 Terra70.3
Claude 3.7 Sonnet67.1
Sonnet 4.666.9
Opus 4.666.7
Claude Fable 566.5
DeepSeek-V3.2-Exp66.3
GLM-5.266.3
Opus 4.765.9
Grok 365.9
Opus 4.865.7
Opus 4.564.3
GLM-5.164.3
Opus 464.1
DeepSeek-V364.1
Claude 3.5 Sonnet (Oct 2024)63.4
Sonnet 463.2
Opus 4.162.9
Sonnet 4.562.8
DeepSeek-V4-Flash62.3
GLM-562.3
Codestral 25.0860.6
GLM 4.660.0
Qwen3-Coder-Next59.1
Kimi K357.1
GPT-4o55.5
GPT-5.6 Luna55.4
GPT-5.6 Luna Pro54.0
GPT-5.3-Codex53.4
Devstral 253.1
Kimi K2.652.8
Gemma 4 31B IT52.1
Gemini 3.5 Flash51.8
GPT-5.2 Instant51.8
GPT-5 Chat (2025-08-07)51.5
Loading Atlas data…