Atlas

Benchmarks

← All benchmarks

GSM8K

Math · 2021-10-27

A grade-school math word problem benchmark focused on multi-step arithmetic reasoning and exact-match solutions.

Top models (higher is better)

ModelScore
Claude 3.5 Sonnet (June 2024)96.4
Opus 395.0
DeepSeek-Coder-V2-Instruct94.5
Qwen2.5 Coder 32B Instruct93.0
Claude 3 Sonnet92.3
GPT-492.0
GPT-4o Mini91.3
Gemini 1.5 Pro90.8
GPT-4 061390.0
Claude 3 Haiku88.9
Phi-3.5-MoE-instruct88.7
Qwen2.5-Coder-14B88.7
DeepSeek-Coder-V2-Lite-Instruct87.6
Claude Instant 1.286.7
Qwen2.5-Coder-7B-Instruct86.7
Phi-3.5 Mini Instruct86.2
Gemma 2 9B84.9
Mistral NeMo Base 240784.2
Qwen2.5 Coder 7B83.9
Llama 3.1 8B Instruct82.4
Claude Instant80.9
PaLM 2-L80.0
Qwen2.5-Coder-3B75.7
Mixtral 8x7B74.4
Stable Beluga 269.6
Yi-34B67.2
DeepSeek-Coder-V2-Lite-Base67.1
Qwen2.5 Coder 1.5B65.8
Grok 162.9
StarCoder2-15B57.7
PaLM 540B56.5
Falcon 180B54.4
Falcon2-11B53.8
GPT-3.5 Turbo 030153.1
Qwen 7B51.7
Gemma 7B46.4
Nemotron-4 15B46.0
Llama 2 34B42.2
CodeQwen1.5 7B37.7
DeepSeek-Coder-Base 33B35.4
Loading Atlas data…