Atlas

Benchmarks

← All benchmarks

Continual Learning Bench Average Gain

Agents · 2026-05-04

CL-Bench normalizes each task's reward gain over the same system's stateless baseline against its published reference ceiling, then averages the six normalized gains equally.

Top models (higher is better)

ModelScore
Sonnet 4.60.2
GPT-5.40.2
Opus 4.70.2
Gemini 3 Flash Preview0.2
Gemini 3.1 Pro Preview0.1
Loading Atlas data…