Atlas

Benchmarks

← All benchmarks

MetagenomicsBench V0

Science

MetagenomicsBench component of benchmarks.bio's original V0 results export, on 100 evaluations. Published pass rates distinguish model and harness. Retained diagnostically because overlapping runs also appear at greater precision in the current V1 export.

Top models (higher is better)

ModelScore
Claude Opus 554.0
Opus 4.850.0
GPT-6 Astra49.7
Sonnet 542.3
Grok 4.642.0
DeepSeek V4.1 Flash38.0
Opus 4.735.3
GPT-5.6 Sol31.3
Kimi K327.3
GPT-5.6 Luna24.3
Loading Atlas data…