Atlas

Benchmarks

← All benchmarks

SpatialBench-Long V1 Verified

Science

SpatialBench-Long V1 verified 22-task subset of the 24-task long-horizon spatial-transcriptomics release. Published per-model and per-harness pass rate.

Top models (higher is better)

ModelScore
GPT-5.6 Terra37.9
GPT-5.6 Sol36.4
DeepSeek V4.1 Flash34.9
GPT-6 Astra34.9
GPT-5.533.3
Claude Opus 531.8
Gemini 3.5 Flash30.3
Grok 4.630.3
Opus 4.828.8
Kimi K328.8
GPT-5.6 Luna27.3
Sonnet 524.2
Grok 4.522.7
Loading Atlas data…