Atlas

Benchmarks

← All benchmarks

ScBench-Long V1 Verified

Science

ScBench-Long V1 verified merged 22-task cohort, distinct from the original 21-task release. Published per-model and per-harness pass rate on extended single-cell analysis workflows.

Top models (higher is better)

ModelScore
GPT-5.6 Sol43.9
GPT-6 Astra42.4
Claude Opus 536.4
GPT-5.6 Terra34.9
GPT-5.6 Luna28.8
DeepSeek V4.1 Flash27.3
Grok 4.627.3
Opus 4.825.8
Grok 4.525.8
Gemini 3.5 Flash24.2
GPT-5.522.7
Sonnet 519.7
Kimi K316.7
Opus 4.710.6
Loading Atlas data…