Atlas

Benchmarks

← All benchmarks

DrugDiscoveryBench

Science · 2026-06-30

DrugDiscoveryBench evaluates coding agents on 82 expert-curated, verifiable multi-step tasks spanning target identification and validation, hit identification, hit-to-lead analysis, and lead optimization in an adapted Biomni tool environment. The primary metric is pass rate: a task passes only when its final-answer outcome score is 100.

Top models (higher is better)

ModelScore
GPT-5.551.6
Sonnet 550.0
Gemini 3.5 Flash50.0
Opus 4.846.8
Gemini 3.1 Pro Preview41.9
GLM-5.236.2
Kimi K2.7 Code35.3
DeepSeek-V4-Pro31.7
Sonnet 4.631.3
GPT-5.2-Codex29.3
Qwen3.7-Max29.3
Opus 4.627.7
MiniMax M322.8
Loading Atlas data…