Atlas

Benchmarks

← All benchmarks

CL-Bench Database Exploration Reward

Agents · 2026-05-04

Mean cumulative reward across 40 sequential instances of the CL-Bench database exploration task.

Top models (higher is better)

ModelScore
Sonnet 4.622.1
GPT-5.417.2
Opus 4.715.7
Gemini 3 Flash Preview15.0
Gemini 3.1 Pro Preview11.6
Loading Atlas data…