DrugDiscoveryBench
Science · 2026-06-30
DrugDiscoveryBench evaluates coding agents on 82 expert-curated, verifiable multi-step tasks spanning target identification and validation, hit identification, hit-to-lead analysis, and lead optimization in an adapted Biomni tool environment. The primary metric is pass rate: a task passes only when its final-answer outcome score is 100.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 51.6 |
| Sonnet 5 | 50.0 |
| Gemini 3.5 Flash | 50.0 |
| Opus 4.8 | 46.8 |
| Gemini 3.1 Pro Preview | 41.9 |
| GLM-5.2 | 36.2 |
| Kimi K2.7 Code | 35.3 |
| DeepSeek-V4-Pro | 31.7 |
| Sonnet 4.6 | 31.3 |
| GPT-5.2-Codex | 29.3 |
| Qwen3.7-Max | 29.3 |
| Opus 4.6 | 27.7 |
| MiniMax M3 | 22.8 |