Atlas

Benchmarks

← All benchmarks

PG-LLM-50

Science · 2026-07-27

PG-LLM-50 evaluates whether general-purpose language models can rank protein variants by experimental fitness from sequence and assay context alone. Across 217 ProteinGym substitution assays, each text-only prompt presents a wild-type sequence, an assay description, and 50 shuffled mutant sequences without measured fitness, examples, alignments, structures, retrieval, or tools. The headline score averages three fixed-draw nested-macro Spearman correlations, weighting assays within proteins and proteins within five functional categories before giving the categories equal weight. Each result uses that model's successfully scored coverage, so assay denominators can differ and cross-model ranks can change on shared-assay subsets.

Top models (higher is better)

ModelScore
Claude Opus 50.4
GPT-5.6 Sol0.4
Opus 4.80.4
Gemini 3.5 Flash0.3
GPT-5.50.3
Gemini 3.1 Pro Preview0.3
Gemini 3.6 Flash0.3
Kimi K30.3
Opus 4.70.3
GPT-5.4 Mini0.2
GPT-5.4 Nano0.2
GLM-5.20.2
Gemini 3.1 Flash-Lite0.2
Loading Atlas data…