Atlas

Benchmarks

← All benchmarks

GeneBench-Pro

Science · 2026-06-30

GeneBench-Pro evaluates AI agents on realistic multi-stage scientific analyses in genomics, quantitative biology, and translational biomedicine. The benchmark contains 129 evaluations across 10 primary domains and 21 terminal subdomains, requiring agents to identify and execute correct analysis workflows across dependent inferential decisions.

Top models (higher is better)

ModelScore
GPT-5.6 Sol Pro31.5
GPT-5.6 Sol28.7
GPT-5.6 Terra Pro28.5
GPT-5.6 Luna Pro23.6
GPT-5.6 Terra23.3
GPT-5.5 Pro20.5
GPT-5.6 Luna16.5
GPT-5.4 Pro16.3
Opus 4.816.0
GPT-5.512.0
GPT-5.48.9
GPT-5.2 Pro8.5
Gemini 3.5 Flash8.1
GPT-5.24.9
GLM-5.24.6
Kimi K2.64.4
Qwen3.7-Max4.0
Gemini 3.1 Pro Preview3.1
DeepSeek-V4-Flash2.4
DeepSeek-V4-Pro2.4
Kimi K2.7 Code2.3
Qwen3.7-Plus2.3
MiMo-V2.5-Pro2.0
Grok 4.31.5
MiMo-V2.51.2
GLM-5.11.2
Hy3 preview0.9
MiniMax M30.9
MiniMax M2.70.6
Loading Atlas data…