Atlas

Benchmarks

← All benchmarks

ZeroEval GSM8K

Math · 2024-08-01

GSM8K evaluated with the ZeroEval zero-shot prompt format, greedy decoding, and answer parser.

Top models (higher is better)

ModelScore
Claude 3.5 Sonnet (Oct 2024)96.7
GPT-4o (2024-08-06)96.2
o1-mini96.1
Llama 3.1 405B Instruct96.0
Claude 3.5 Sonnet (June 2024)95.6
Opus 395.6
Mistral Large 2 (Instruct 2407)95.5
GPT-4o95.4
Gemini 1.5 Pro Experimental 080195.0
Haiku 3.594.5
GPT-4o Mini94.2
Llama 3.1 70B Instruct94.2
DeepSeek-V2-Chat-062893.9
DeepSeek-Coder-V2-Instruct93.8
Gemini 1.5 Pro93.4
Llama 3 70B Instruct93.0
Qwen2 72B Instruct92.7
DeepSeek-V2.592.5
DeepSeek-Coder-V2-Instruct-072491.5
Claude 3 Sonnet91.5
Gemini 1.5 Flash91.4
Gemma 2 27B IT90.2
Claude 3 Haiku88.8
Gemma 2 9B IT87.4
Reka Core 20240501 (ZeroEval label)87.4
Athene-70B86.7
Yi 1.5 34B Chat84.1
Llama 3.1 8B Instruct84.0
Mistral NeMo Instruct 240782.8
Yi-Large82.6
Phi-3.5 Mini Instruct82.0
GPT-3.5 Turbo 012580.4
C4AI Command R+80.1
Qwen2 7B Instruct80.1
Llama 3 8B Instruct78.5
Yi 1.5 9B Chat76.4
Phi-3-mini-4k-instruct75.5
Reka Flash (2024-02-26)74.7
Mixtral 8x7B Instruct70.1
Command R53.0
Loading Atlas data…