SLDBench Task-Mean R²
Science · 2025-12-07
This row reports the mean R² averaged across SLDBench task leaderboards using the official site aggregation, including padded invalid runs.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5 | 0.7 |
| O4 Mini | 0.7 |
| Gemini 3 Pro Preview | 0.6 |
| Sonnet 4.5 | 0.6 |
| Haiku 4.5 | 0.5 |
| Gemini 2.5 Flash | 0.5 |
| GPT-5.2 | 0.2 |
| GPT-4.1 | 0.1 |
| o3 | 0.0 |
| DeepSeek-V3.2 | 0.0 |
| Gemini 3 Flash Preview | -0.7 |
| GPT-4o | -0.8 |