Atlas

Benchmarks

← All benchmarks

MGSM English

Math · 2022-10-06

English language split of the MGSM multilingual grade-school math benchmark.

Top models (higher is better)

ModelScore
MiniMax M2.1100.0
Opus 4.599.6
GLM 4.699.2
Mistral Medium 399.2
O4 Mini99.2
Sonnet 498.8
Sonnet 4.598.8
Gemini 3 Flash Preview98.8
Gemini 3 Pro Preview98.8
GPT-5.198.8
GPT-5.298.8
Qwen3 Max Preview98.8
Haiku 4.598.4
DeepSeek-V3-032498.4
GLM-4.798.4
GPT-598.4
GPT-5 Mini98.4
Llama 4 Maverick Instruct98.4
Mistral Large 3 675B Instruct 251298.4
o3-mini98.4
Qwen3 Max (rolling alias)98.4
Claude 3.5 Sonnet (Oct 2024)98.0
Claude 3.7 Sonnet98.0
Opus 498.0
DeepSeek-R198.0
DeepSeek-V3.298.0
Gemini 2.5 Flash Preview (09-2025)98.0
gpt-oss-120b98.0
Kimi K2 Instruct98.0
Magistral Small 1.298.0
Command A97.6
Opus 4.197.6
DeepSeek-V397.6
GPT-5 Nano97.6
gpt-oss-20b97.6
O197.6
o397.6
Qwen3-235B-A22B97.6
GPT-4.1 Mini97.2
GPT-4o (2024-08-06)97.2
Loading Atlas data…