Atlas

Benchmarks

← All benchmarks

GSM1k

Math

Scale's 1,000-question held-out GSM8K-style mathematics benchmark, evaluated with a five-shot prompt.

Top models (higher is better)

ModelScore
Claude 3.5 Sonnet (June 2024)96.6
GPT-4o (2024-08-06)95.7
Llama 3.1 405B Instruct95.6
Opus 395.2
GPT-4 0125 Preview95.1
GPT-4o94.8
Gemini 1.5 Pro Experimental 082794.7
Mistral Large 2 (Instruct 2407)93.9
Claude 3 Sonnet93.3
Gemini 1.5 Pro Preview 051492.3
Gemini 1.5 Pro Preview 040990.5
Gemini 1.5 Flash Preview 051490.1
Llama 3 70B Instruct90.1
Mistral Large 1.087.5
Gemini 1.0 Pro79.8
Code Llama Instruct 34B37.5
Loading Atlas data…