Atlas

Benchmarks

← All benchmarks

ZeroEval ZebraLogic

Math · 2025-02-03

ZebraLogic grid puzzles evaluated with the ZeroEval zero-shot prompt format, greedy decoding, and puzzle-level accuracy metric.

Top models (higher is better)

ModelScore
Grok 3 Mini92.6
o3-mini91.7
O181.0
DeepSeek-R178.7
o1 Preview71.4
Grok 357.9
o1-mini52.6
DeepSeek-V342.1
Claude 3.5 Sonnet (Oct 2024)36.2
Claude 3.5 Sonnet (June 2024)33.4
Llama 3.1 405B Instruct32.6
GPT-4o (2024-08-06)31.7
Gemini 1.5 Pro Experimental 082730.5
ChatGPT-4o Latest (September 2024 benchmark entry)29.9
Mistral Large 2 (Instruct 2407)29.0
GPT-4 Turbo28.4
GPT-4o28.2
Grok 2 121227.7
GPT-427.1
Opus 327.0
Qwen2.5 72B Instruct26.6
Qwen2.5 Instruct 32B26.1
Gemini 1.5 Pro Experimental 080125.2
Gemini 1.5 Flash Experimental 082725.0
Llama 3.1 70B Instruct24.9
DeepSeek-V2-Chat-062822.7
DeepSeek-V2.522.1
Qwen2 72B Instruct21.4
DeepSeek-Coder-V2-Instruct21.1
DeepSeek-Coder-V2-Instruct-072420.5
GPT-4o Mini20.1
Gemini 1.5 Flash19.4
Gemini 1.5 Pro19.4
Yi-Large18.9
Haiku 3.518.7
Claude 3 Sonnet18.7
Llama 3 70B Instruct16.8
Athene-70B16.7
Gemma 2 27B IT16.3
Claude 3 Haiku14.3
Loading Atlas data…