GeneBench-Pro
Science · 2026-06-30
GeneBench-Pro evaluates AI agents on realistic multi-stage scientific analyses in genomics, quantitative biology, and translational biomedicine. The benchmark contains 129 evaluations across 10 primary domains and 21 terminal subdomains, requiring agents to identify and execute correct analysis workflows across dependent inferential decisions.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol Pro | 31.5 |
| GPT-5.6 Sol | 28.7 |
| GPT-5.6 Terra Pro | 28.5 |
| GPT-5.6 Luna Pro | 23.6 |
| GPT-5.6 Terra | 23.3 |
| GPT-5.5 Pro | 20.5 |
| GPT-5.6 Luna | 16.5 |
| GPT-5.4 Pro | 16.3 |
| Opus 4.8 | 16.0 |
| GPT-5.5 | 12.0 |
| GPT-5.4 | 8.9 |
| GPT-5.2 Pro | 8.5 |
| Gemini 3.5 Flash | 8.1 |
| GPT-5.2 | 4.9 |
| GLM-5.2 | 4.6 |
| Kimi K2.6 | 4.4 |
| Qwen3.7-Max | 4.0 |
| Gemini 3.1 Pro Preview | 3.1 |
| DeepSeek-V4-Flash | 2.4 |
| DeepSeek-V4-Pro | 2.4 |
| Kimi K2.7 Code | 2.3 |
| Qwen3.7-Plus | 2.3 |
| MiMo-V2.5-Pro | 2.0 |
| Grok 4.3 | 1.5 |
| MiMo-V2.5 | 1.2 |
| GLM-5.1 | 1.2 |
| Hy3 preview | 0.9 |
| MiniMax M3 | 0.9 |
| MiniMax M2.7 | 0.6 |