Atlas

Benchmarks

← All benchmarks

CL-Bench Codebase Adaptation Reward

Agents · 2026-05-04

Mean cumulative reward across 19 sequential instances of the CL-Bench codebase adaptation task.

Top models (higher is better)

ModelScore
GPT-5.411.6
Opus 4.710.4
Sonnet 4.69.8
Gemini 3 Flash Preview7.4
Gemini 3.1 Pro Preview7.1
Loading Atlas data…