Atlas

Benchmarks

← All benchmarks

HellaSwag (Unspecified Scoring Protocol)

General QA · 2019-05-19

HellaSwag results for which the reporting source does not establish whether raw or length-normalized multiple-choice scoring was used. Values are retained as reported but quarantined from protocol-specific comparisons.

Top models (higher is better)

ModelScore
Opus 395.4
GPT-495.3
GPT-4 32K (0314)95.3
Gemini 1.5 Pro93.3
Llama 3.1 405B89.2
Falcon 180B89.0
Claude 3 Sonnet89.0
DeepSeek-V388.9
Gemini 1.0 Ultra87.8
DeepSeek-V287.1
PaLM 2-L86.8
Gemini 1.5 Flash86.5
Claude 3 Haiku85.9
Llama 2 70B85.3
Qwen2.5 72B84.8
Gemini 1.0 Pro84.7
LLaMA 65B84.2
Stable Beluga 284.1
PaLM 2-M84.0
LLaMA 33B82.8
Nemotron-4 15B82.4
Phi-3 Medium 128K Instruct82.4
PaLM 2-S82.0
Mistral 7B v0.1 Base81.0
Llama 2 13B80.7
Qwen2.5-Coder-14B80.2
Gopher (280B)79.2
LLaMA-13B79.2
OPT-175B79.1
InternLM 20B78.1
GPT-3 Davinci77.5
Phi-3-Small-8K-Instruct77.0
Qwen2.5 Coder 7B76.8
Phi-3-mini-4k-instruct76.7
Yi-9B76.4
Yi-6B74.4
XGen-7B-8K-Base74.2
OpenLLaMA 7B71.8
Gemma 2B71.4
Qwen2.5-Coder-3B70.9
Loading Atlas data…