Atlas

Benchmarks

← All benchmarks

SpatialBench V1 Verified

Science

SpatialBench V1 verified subset of 115 spatial-transcriptomics tasks from the original 159-task release. Published per-model and per-harness pass rate.

Top models (higher is better)

ModelScore
Grok 4.677.4
DeepSeek V4.1 Flash75.4
GPT-5.6 Sol74.5
Claude Opus 571.0
Grok 4.571.0
GPT-5.6 Terra70.4
GPT-6 Astra70.1
Sonnet 569.3
Opus 4.869.0
Gemini 3.5 Flash67.0
GPT-5.565.2
Opus 4.764.1
Kimi K361.5
GPT-5.6 Luna60.6
Loading Atlas data…