Atlas

Benchmarks

← All benchmarks

BioMysteryBench — Hard (Vals)

Science · 2026-09-22

The 17 human-difficult bioinformatics tasks, with three Terminus 2 trials per task.

Top models (higher is better)

ModelScore
Claude Opus 5.547.1
Claude Opus 543.1
Hy4 Preview39.2
GPT-6 Astra37.3
GPT-6 Sol35.3
Kimi K335.3
GPT-5.6 Sol29.4
Grok 4.629.4
Grok 4.729.4
GPT-5.6 Luna27.5
Muse Spark 1.225.5
DeepSeek-V4-Flash-073123.5
DeepSeek V4.1 Flash21.6
Gemini 3.6 Flash21.6
Gemini 3.8 Flash21.6
GPT-6 Luna17.6
Loading Atlas data…