MLE-Bench Partial 30 - Average Position Score
Agents · 2024-10-09
Google’s MLE-Bench Partial 30 evaluation covers the canonical 30-competition subset and reports Average Position Score over two independent runs per problem. Position Score is the model submission’s percentile-like rank on each historical Kaggle private leaderboard, with 1.0 for first place and 0 for a missing or failed submission.
Top models (higher is better)
| Model | Score |
|---|---|
| Sonnet 5 | 66.9 |
| Gemini 3.6 Flash | 63.9 |
| Gemini 3.5 Flash | 49.7 |
| GPT-5.6 Luna | 47.6 |
| Grok 4.5 | 43.2 |
| Gemini 3.1 Pro Preview | 42.6 |
| Gemini 3.5 Flash-Lite | 39.2 |
| Gemini 3.1 Flash-Lite | 22.0 |