Atlas

Benchmarks

← All benchmarks

MGSM Spanish

Math · 2022-10-06

Spanish language split of the MGSM multilingual grade-school math benchmark.

Top models (higher is better)

ModelScore
Opus 4.597.2
Sonnet 4.596.8
Opus 4.196.4
Sonnet 496.4
Gemini 3 Flash Preview96.4
Gemini 3 Pro Preview96.4
O4 Mini96.4
Claude 3.5 Sonnet (Oct 2024)96.0
Opus 496.0
DeepSeek-V396.0
Claude 3.7 Sonnet95.6
DeepSeek-V3-032495.6
gpt-oss-120b95.2
Kimi K2 Instruct95.2
Qwen3-235B-A22B95.2
DeepSeek-R194.8
GPT-5 Mini94.8
Grok 394.8
Qwen3 Max Preview94.8
GPT-5.294.4
Mistral Medium 394.4
GPT-4.1 Mini94.0
Mistral Large 3 675B Instruct 251294.0
DeepSeek-V3.293.6
Kimi K2 Thinking93.6
Llama 4 Maverick Instruct93.6
Gemini 2.5 Flash-Lite Preview (09-2025)93.2
GPT-5 Nano93.2
Grok 293.2
Grok 493.2
Llama 3.3 70B Instruct93.2
Qwen3 Max (rolling alias)93.2
Command A92.8
Gemini 2.0 Flash 00192.8
Magistral Small 1.292.8
Gemini 1.5 Pro 00292.4
GLM-4.592.4
Grok 3 Mini92.4
Magistral Medium 1.292.4
Haiku 4.592.0
Loading Atlas data…