CL-Bench Database Exploration Gain
Agents · 2026-05-04
Mean cumulative gain over the stateless baseline across 40 sequential instances of the CL-Bench database exploration task.
Top models (higher is better)
| Model | Score |
|---|---|
| Sonnet 4.6 | 13.9 |
| GPT-5.4 | 12.9 |
| Gemini 3 Flash Preview | 11.5 |
| Opus 4.7 | 9.6 |
| Gemini 3.1 Pro Preview | 6.8 |