Atlas

Benchmarks

← All benchmarks

scBench

Science · 2026-02-06

scBench contains 195 verifiable agent tasks drawn from practical single-cell RNA-seq workflows. Deterministic graders report the percentage of tasks passed.

Top models (higher is better)

ModelScore
Claude Opus 560.6
Claude Mythos 559.3
Opus 4.858.2
GPT-5.558.0
GPT-5.457.4
Gemini 3.5 Flash56.9
Sonnet 556.2
Opus 4.755.2
Gemini 3.1 Pro Preview53.9
Opus 4.652.6
GPT-5.252.3
Sonnet 4.650.3
Opus 4.547.2
Grok 4.2044.4
Grok 4.344.3
GPT-5.138.8
Sonnet 4.533.2
Grok 4.1 Fast30.3
Gemini 2.5 Pro23.6
Loading Atlas data…