Atlas

Benchmarks

← All benchmarks

ZeroEval

Indexes · 2024-08-01

ZeroEval is Allen AI's zero-shot evaluation suite for language models. Its headline score is the official average over MMLU-Redux, ZebraLogic, CRUX, and MATH Level 5 using the ZeroEval prompting and answer-parsing setup.

Top models (higher is better)

ModelScore
o1 Preview86.1
o1-mini80.6
Claude 3.5 Sonnet (Oct 2024)67.1
Gemini 1.5 Pro Experimental 082766.1
GPT-4o (2024-08-06)65.6
ChatGPT-4o Latest (September 2024 benchmark entry)64.6
GPT-4o64.3
Claude 3.5 Sonnet (June 2024)63.0
Grok 2 121262.8
Qwen2.5 72B Instruct61.6
Llama 3.1 405B Instruct59.8
GPT-4 Turbo59.8
Gemini 1.5 Flash Experimental 082759.0
Mistral Large 2 (Instruct 2407)58.9
GPT-4o Mini57.4
DeepSeek-V2.554.3
Opus 354.2
Llama 3.1 70B Instruct53.8
Haiku 3.553.4
Gemini 1.5 Pro52.5
GPT-452.3
Qwen2 72B Instruct50.1
Gemini 1.5 Flash48.8
Qwen2.5 7B Instruct47.8
Llama 3 70B Instruct44.7
Gemma 2 27B IT44.0
Athene-70B41.2
Reka Core 20240501 (ZeroEval label)39.4
Claude 3 Haiku39.1
Gemma 2 9B IT37.8
GPT-3.5 Turbo 012536.7
Yi 1.5 34B Chat36.6
Phi-3-mini-4k-instruct35.7
Llama 3.1 8B Instruct35.5
Qwen2 7B Instruct34.3
Phi-3.5 Mini Instruct33.7
Yi 1.5 9B Chat33.0
Qwen2.5 3B Instruct31.9
Llama 3 8B Instruct29.8
Gemma 2 2B IT20.5
Loading Atlas data…