SimpleBench Open-Ended
Classic NLP · 2024-08-20
SimpleBench Open-Ended evaluates the same everyday reasoning tasks without exposing multiple-choice answer options. The official leaderboard reports accuracy averaged across five runs (AVG@5).
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 3.1 Pro Preview | 64.8 |
| Opus 4.7 | 62.9 |
| Opus 4.6 | 59.7 |