Atlas

Benchmarks

← All benchmarks

Ruby LLM Benchmarks — Program Fixer

Code · 2025-08-01

Ruby LLM Benchmarks Program Fixer evaluates one-shot repairs of four broken Ruby programs: a calendar, parking garage, school library, and vending machine. Each task score is 90% test-suite success plus 10% RuboCop quality; the headline score is the equal-weight mean of the four task scores, with a missing task contributing zero.

Top models (higher is better)

ModelScore
Claude Opus 579.8
GPT-5.6 Terra Pro78.1
GPT-5.577.6
Claude Fable 577.2
Opus 4.777.1
Opus 4.676.4
Sonnet 4.676.3
GLM-575.6
GPT-5.6 Sol Pro75.3
Opus 4.875.2
GPT-5.6 Sol75.2
GPT-5.474.8
GLM-5.274.2
Kimi K2.674.2
GPT-5.6 Terra74.0
GPT-5.6 Luna73.3
GPT-5.6 Luna Pro73.1
Sonnet 472.6
DeepSeek-V4-Pro72.5
Sonnet 4.572.3
GPT-5.2 Instant72.1
Opus 4.571.6
Opus 4.171.3
Gemma 4 26B A4B IT71.0
Gemma 4 31B IT70.9
Opus 470.8
GPT-4.170.8
Sonnet 570.3
GPT-5.3 Instant70.2
Grok 4.570.0
Horizon Beta69.8
GPT-5.1 Instant69.8
Kimi K369.6
GPT-5.1-Codex69.4
DeepSeek-V3.2-Speciale69.3
Kimi K2 Thinking69.2
DeepSeek-V4-Flash68.8
GPT-568.8
ChatGPT-4o Latest (source-unspecified snapshot)68.7
GPT-5.3-Codex68.6
Loading Atlas data…