Atlas

Benchmarks

← All benchmarks

SLDBench Task-Mean R²

Science · 2025-12-07

This row reports the mean R² averaged across SLDBench task leaderboards using the official site aggregation, including padded invalid runs.

Top models (higher is better)

ModelScore
GPT-50.7
O4 Mini0.7
Gemini 3 Pro Preview0.6
Sonnet 4.50.6
Haiku 4.50.5
Gemini 2.5 Flash0.5
GPT-5.20.2
GPT-4.10.1
o30.0
DeepSeek-V3.20.0
Gemini 3 Flash Preview-0.7
GPT-4o-0.8
Loading Atlas data…