Atlas

Benchmarks

← All benchmarks

Code Migration - COBOL

Code · 2026-06-16

The COBOL split of Code Migration. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Opus 4.770.0
Claude Opus 570.0
GLM-5.270.0
GPT-5.570.0
GPT-5.6 Luna70.0
GPT-5.6 Sol70.0
Grok 4.570.0
Opus 4.868.6
Sonnet 4.668.6
Sonnet 568.6
GPT-5.468.6
Muse Spark 1.168.6
GPT-5.6 Terra65.8
Gemini 3.6 Flash65.7
Kimi K2.7 Code61.3
DeepSeek-V4-Pro58.4
Kimi K2.657.1
Gemini 3.5 Flash55.5
GLM-5.146.6
MiMo-V2.5-Pro44.3
MiniMax M343.9
Claude Fable 540.0
Gemini 3.1 Pro Preview36.8
MiMo-V2.527.3
GPT-5.4 Mini27.1
Inkling26.1
Gemini 3.5 Flash-Lite7.0
MiniMax M2.74.3
Gemini 3.1 Flash-Lite Preview1.3
Qwen3.6 Plus (2026-04-02)1.3
Haiku 4.50.0
Gemini 3 Flash Preview0.0
Grok 4.30.0
Loading Atlas data…