Atlas

Benchmarks

← All benchmarks

ZeroEval MATH Level 5

Math · 2024-08-01

MATH Level 5 evaluated with the ZeroEval zero-shot prompt format, greedy decoding, and answer parser.

Top models (higher is better)

ModelScore
o1-mini89.3
o1 Preview84.5
Gemini 1.5 Pro Experimental 082768.1
Grok 2 121260.9
Qwen2.5 72B Instruct60.2
Qwen2.5 Instruct 32B59.5
Claude 3.5 Sonnet (Oct 2024)59.4
GPT-4o (2024-08-06)55.3
GPT-4o54.8
Gemini 1.5 Flash Experimental 082754.5
ChatGPT-4o Latest (September 2024 benchmark entry)53.1
GPT-4o Mini52.1
Claude 3.5 Sonnet (June 2024)51.9
Qwen2.5 7B Instruct51.5
Llama 3.1 405B Instruct49.8
Mistral Large 2 (Instruct 2407)48.5
Haiku 3.546.5
GPT-4 Turbo46.5
DeepSeek-V2.544.7
Llama 3.1 70B Instruct43.1
Gemini 1.5 Pro39.8
Qwen2 72B Instruct38.3
Opus 336.9
Gemini 1.5 Flash34.8
Gemma 2 27B IT26.6
GPT-426.1
Qwen2.5 3B Instruct25.5
Llama 3 70B Instruct25.1
Qwen2 7B Instruct23.9
Llama 3.1 8B Instruct22.2
Reka Core 20240501 (ZeroEval label)21.9
Athene-70B20.7
Yi 1.5 9B Chat20.0
Gemma 2 9B IT19.4
Phi-3.5 Mini Instruct18.7
Yi 1.5 34B Chat18.2
Phi-3-mini-4k-instruct16.2
Claude 3 Haiku15.1
GPT-3.5 Turbo 012513.7
Llama 3 8B Instruct7.9
Loading Atlas data…