Ruby LLM Program Fixer — Parking Garage
Code · 2025-08-01
The Parking Garage Program Fixer task asks a model to repair a broken Ruby parking-garage implementation against 39 tests. Its composite is 90% test success plus 10% RuboCop quality, with each offense subtracting two quality points down to zero.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 85.0 |
| GPT-5.6 Sol | 80.1 |
| GPT-5.6 Sol Pro | 79.9 |
| GPT-5.4 | 77.0 |
| GPT-5.6 Terra Pro | 77.0 |
| Claude Opus 5 | 72.0 |
| GPT-5.6 Terra | 70.3 |
| Claude 3.7 Sonnet | 67.1 |
| Sonnet 4.6 | 66.9 |
| Opus 4.6 | 66.7 |
| Claude Fable 5 | 66.5 |
| DeepSeek-V3.2-Exp | 66.3 |
| GLM-5.2 | 66.3 |
| Opus 4.7 | 65.9 |
| Grok 3 | 65.9 |
| Opus 4.8 | 65.7 |
| Opus 4.5 | 64.3 |
| GLM-5.1 | 64.3 |
| Opus 4 | 64.1 |
| DeepSeek-V3 | 64.1 |
| Claude 3.5 Sonnet (Oct 2024) | 63.4 |
| Sonnet 4 | 63.2 |
| Opus 4.1 | 62.9 |
| Sonnet 4.5 | 62.8 |
| DeepSeek-V4-Flash | 62.3 |
| GLM-5 | 62.3 |
| Codestral 25.08 | 60.6 |
| GLM 4.6 | 60.0 |
| Qwen3-Coder-Next | 59.1 |
| Kimi K3 | 57.1 |
| GPT-4o | 55.5 |
| GPT-5.6 Luna | 55.4 |
| GPT-5.6 Luna Pro | 54.0 |
| GPT-5.3-Codex | 53.4 |
| Devstral 2 | 53.1 |
| Kimi K2.6 | 52.8 |
| Gemma 4 31B IT | 52.1 |
| Gemini 3.5 Flash | 51.8 |
| GPT-5.2 Instant | 51.8 |
| GPT-5 Chat (2025-08-07) | 51.5 |