CL-Bench Cohort Studies Reward
Agents · 2026-05-04
Mean cumulative reward across 20 sequential instances of the CL-Bench cohort studies task.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.4 | -0.1 |
| Opus 4.7 | -0.1 |
| Gemini 3 Flash Preview | -0.5 |
| Sonnet 4.6 | -0.8 |
| Gemini 3.1 Pro Preview | -1.0 |