Atlas

Benchmarks

← All benchmarks

MLS Bench Lite

Science · 2026-05-05

MLS-Bench-Lite is the official 30-task subset spanning all 12 MLS-Bench research domains and evaluates agents on generalizable machine-learning science contributions.

Top models (higher is better)

ModelScore
Claude Fable 549.9
Kimi K348.3
GPT-5.6 Sol46.2
Opus 4.842.8
GLM-5.240.4
GPT-5.535.5
Loading Atlas data…