Atlas

Benchmarks

← All benchmarks

Code Migration - CLI

Code · 2026-06-16

The CLI split of Code Migration. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Fable 560.1
Claude Opus 553.3
GPT-5.6 Sol47.2
Opus 4.840.1
GPT-5.536.9
Sonnet 536.3
GPT-5.6 Luna36.1
Opus 4.735.2
Sonnet 4.630.3
GLM-5.227.2
GPT-5.6 Terra26.6
Grok 4.525.5
GPT-5.423.8
Gemini 3.6 Flash19.3
GLM-5.118.8
Muse Spark 1.118.6
Kimi K2.618.0
Gemini 3.5 Flash17.2
DeepSeek-V4-Pro15.5
Qwen3.6 Plus (2026-04-02)14.4
MiMo-V2.5-Pro14.0
Kimi K2.7 Code13.4
MiniMax M311.9
Gemini 3.1 Pro Preview10.8
MiniMax M2.710.2
Haiku 4.510.1
MiMo-V2.59.9
Grok 4.39.1
Gemini 3 Flash Preview8.5
GPT-5.4 Mini8.2
Inkling7.0
Gemini 3.5 Flash-Lite5.9
Gemini 3.1 Flash-Lite Preview5.7
Loading Atlas data…