Atlas

Benchmarks

← All benchmarks

VariantBench

Science · 2026-07-16

VariantBench evaluates agents on 118 real-world genetic-variant discovery, quality-control, statistical genetics, population genomics, and clinical interpretation tasks. Scores report the percentage of task runs passed.

Top models (higher is better)

ModelScore
Opus 4.842.1
GPT-5.6 Sol42.1
Sonnet 538.7
GPT-5.6 Terra38.7
Grok 4.538.1
Opus 4.737.0
Gemini 3.5 Flash35.3
Gemini 3.1 Pro Preview34.5
Kimi K332.8
GPT-5.530.2
GPT-5.6 Luna29.9
Grok 4.320.3
Loading Atlas data…