Atlas

Benchmarks

← All benchmarks

BioMysteryBench — Human-solvable (Vals)

Science · 2026-09-22

The 73 human-solvable bioinformatics tasks, with three Terminus 2 trials per task.

Top models (higher is better)

ModelScore
GPT-6 Astra89.0
Claude Opus 587.7
Claude Opus 5.586.8
GPT-6 Sol84.0
Grok 4.682.2
GPT-5.6 Sol80.8
Kimi K379.9
DeepSeek V4.1 Flash78.5
Grok 4.778.5
Hy4 Preview76.3
DeepSeek-V4-Flash-073174.0
Muse Spark 1.274.0
Gemini 3.8 Flash71.7
GPT-6 Luna71.7
GPT-5.6 Luna69.4
Gemini 3.6 Flash67.1
Loading Atlas data…