CL-Bench Codebase Adaptation Reward
Agents · 2026-05-04
Mean cumulative reward across 19 sequential instances of the CL-Bench codebase adaptation task.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.4 | 11.6 |
| Opus 4.7 | 10.4 |
| Sonnet 4.6 | 9.8 |
| Gemini 3 Flash Preview | 7.4 |
| Gemini 3.1 Pro Preview | 7.1 |