Atlas

Benchmarks

← All benchmarks

Apex Shortlist (MathArena IRT estimate)

Math

MathArena estimated accuracy combines measured answers with item-response-theory predictions for unanswered questions. Retained as a diagnostic, separate from observed accuracy and excluded from the capability fit.

Top models (higher is better)

ModelScore
Qwen3.5-4B32.5
Qwen3.5-2B8.0
Loading Atlas data…