Atlas

Benchmarks

← All benchmarks

ScBench-Long V2

Science

ScBench-Long V2 clean-v1 reruns on 22 long-horizon single-cell analysis evaluations. Macro-mean pass rate; the aggregate site's weight of 21 is not the task count.

Top models (higher is better)

ModelScore
GPT-5.6 Sol34.9
Claude Opus 530.3
GPT-6 Astra25.8
Opus 4.824.2
Gemini 3.8 Flash24.2
GPT-5.524.2
GPT-6 Sol24.2
Gemini 3.7 Flash19.7
Grok 4.619.7
Sonnet 515.2
GPT-5.6 Luna15.2
Opus 4.712.1
Sonnet 4.612.1
GPT-6 Luna10.6
Grok 4.79.1
Claude Opus 5.57.6
Loading Atlas data…