Atlas

Benchmarks

← All benchmarks

ScBench-Long

Science · 2026-06-25

Long-horizon single-cell analysis benchmark on real single-cell RNA-seq datasets, requiring extended multi-step agentic workflows such as quality control, clustering, cell-type annotation, and downstream analysis. Higher pass rates indicate better performance.

Top models (higher is better)

ModelScore
GPT-5.6 Sol38.1
Opus 4.825.4
GPT-5.6 Terra23.8
Gemini 3.5 Flash22.2
GPT-5.6 Luna19.1
Gemini 3.1 Pro Preview17.5
GPT-5.515.9
Sonnet 514.3
Grok 4.512.7
Kimi K312.7
Opus 4.711.1
Opus 4.69.5
GLM-5.29.5
GPT-5.49.5
Sonnet 4.67.9
Grok 4.206.3
Kimi K2.64.8
Grok 4.31.6
Loading Atlas data…