CL-Bench Database Exploration Reward
Agents · 2026-05-04
Mean cumulative reward across 40 sequential instances of the CL-Bench database exploration task.
Top models (higher is better)
| Model | Score |
|---|---|
| Sonnet 4.6 | 22.1 |
| GPT-5.4 | 17.2 |
| Opus 4.7 | 15.7 |
| Gemini 3 Flash Preview | 15.0 |
| Gemini 3.1 Pro Preview | 11.6 |