Atlas

Benchmarks

← All benchmarks

BioSecBench-Function V1

Science

BioSecBench-Function V1: 111 deterministic biological-function evaluation tasks on experimental data. Published pass rates distinguish model and agent harness.

Top models (higher is better)

ModelScore
Claude Opus 550.3
GPT-6 Astra48.6
Opus 4.844.6
Grok 4.644.1
Gemini 3.5 Flash42.2
DeepSeek V4.1 Flash38.7
GPT-5.6 Sol38.3
Sonnet 537.2
GPT-5.536.5
Grok 4.535.6
GPT-5.6 Terra33.5
GPT-5.6 Luna26.9
Loading Atlas data…