Atlas

Benchmarks

← All benchmarks

Continual Learning Bench (CL-Bench)

Agents · 2026-05-04

Continual Learning Bench evaluates stateful agents that adapt across 301 sequential instances in six expert-validated task domains. This headline metric normalizes each task's reward against a reference ceiling and the GPT-5.4 stateless baseline, then averages the six normalized task scores.

Top models (higher is better)

ModelScore
Sonnet 4.60.2
GPT-5.40.2
Opus 4.70.2
Gemini 3 Flash Preview0.1
Gemini 3.1 Pro Preview0.0
Loading Atlas data…