Atlas

Benchmarks

← All benchmarks

MLE-Bench Partial 30 - Average Position Score

Agents · 2024-10-09

Google’s MLE-Bench Partial 30 evaluation covers the canonical 30-competition subset and reports Average Position Score over two independent runs per problem. Position Score is the model submission’s percentile-like rank on each historical Kaggle private leaderboard, with 1.0 for first place and 0 for a missing or failed submission.

Top models (higher is better)

ModelScore
Sonnet 566.9
Gemini 3.6 Flash63.9
Gemini 3.5 Flash49.7
GPT-5.6 Luna47.6
Grok 4.543.2
Gemini 3.1 Pro Preview42.6
Gemini 3.5 Flash-Lite39.2
Gemini 3.1 Flash-Lite22.0
Loading Atlas data…