Atlas

Benchmarks

← All benchmarks

SpatialBench-Long V2

Science

SpatialBench-Long V2 clean-v1 reruns on 21 long-horizon spatial-transcriptomics evaluations. Macro-mean pass rate; the aggregate site's weight of 24 is not the task count.

Top models (higher is better)

ModelScore
Claude Opus 541.3
GPT-6 Astra39.7
GPT-5.6 Sol36.5
GPT-5.534.9
Opus 4.833.3
Opus 4.730.2
GPT-6 Sol28.6
Grok 4.627.0
Gemini 3.7 Flash25.4
Sonnet 4.623.8
Sonnet 523.8
Gemini 3.8 Flash23.8
GPT-5.6 Luna23.8
GPT-6 Luna23.8
Grok 4.723.8
Claude Opus 5.519.1
Loading Atlas data…