Atlas

Benchmarks

← All benchmarks

CL-Bench Blind Spectrum Monitoring Reward

Agents · 2026-05-04

Mean cumulative reward across 90 sequential instances of the CL-Bench blind spectrum monitoring task.

Top models (higher is better)

ModelScore
GPT-5.446.2
Sonnet 4.644.3
Opus 4.733.6
Gemini 3 Flash Preview33.0
Gemini 3.1 Pro Preview33.0
Loading Atlas data…