Atlas

Benchmarks

← All benchmarks

SimpleBench Open-Ended

Classic NLP · 2024-08-20

SimpleBench Open-Ended evaluates the same everyday reasoning tasks without exposing multiple-choice answer options. The official leaderboard reports accuracy averaged across five runs (AVG@5).

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview64.8
Opus 4.762.9
Opus 4.659.7
Loading Atlas data…