MetagenomicsBench V0
Science
MetagenomicsBench component of benchmarks.bio's original V0 results export, on 100 evaluations. Published pass rates distinguish model and harness. Retained diagnostically because overlapping runs also appear at greater precision in the current V1 export.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 54.0 |
| Opus 4.8 | 50.0 |
| GPT-6 Astra | 49.7 |
| Sonnet 5 | 42.3 |
| Grok 4.6 | 42.0 |
| DeepSeek V4.1 Flash | 38.0 |
| Opus 4.7 | 35.3 |
| GPT-5.6 Sol | 31.3 |
| Kimi K3 | 27.3 |
| GPT-5.6 Luna | 24.3 |