Atlas

Benchmarks

← All benchmarks

MGSM

Math · 2022-10-06

MGSM is a multilingual grade-school mathematics benchmark built from manually translated GSM8K problems in ten languages. Scores are reported as exact-answer accuracy.

Top models (higher is better)

ModelScore
Opus 4.595.2
Opus 4.194.4
Sonnet 4.594.3
GPT-5.294.0
Gemini 3 Pro Preview93.9
Opus 493.8
O4 Mini93.4
Gemini 3 Flash Preview93.3
Sonnet 493.0
Claude 3.7 Sonnet93.0
GPT-5.193.0
GPT-592.8
Claude 3.5 Sonnet (Oct 2024)92.6
GPT-5 Mini92.6
Qwen3-235B-A22B92.5
Llama 4 Maverick Instruct92.4
DeepSeek-R192.3
Haiku 4.592.1
DeepSeek-V392.1
Qwen3 Max Preview92.1
gpt-oss-120b92.0
Qwen3 Max (rolling alias)91.8
o391.7
DeepSeek-V3-032491.7
Claude 3.5 Sonnet (June 2024)91.6
Grok 391.3
o3-mini91.3
Llama 3.3 70B Instruct91.1
Kimi K2 Instruct90.9
Grok 490.9
DeepSeek-V3.290.9
Grok 4 Fast90.9
Mistral Medium 390.9
GLM-4.590.8
GPT-4o (2024-08-06)90.7
Grok 3 Mini90.4
GPT-4o (2024-11-20)90.4
Kimi K2 Thinking90.1
Gemini 2.5 Flash Preview (09-2025)89.9
GLM 4.689.7
Loading Atlas data…