HellaSwag (Unspecified Scoring Protocol)
General QA · 2019-05-19
HellaSwag results for which the reporting source does not establish whether raw or length-normalized multiple-choice scoring was used. Values are retained as reported but quarantined from protocol-specific comparisons.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 3 | 95.4 |
| GPT-4 | 95.3 |
| GPT-4 32K (0314) | 95.3 |
| Gemini 1.5 Pro | 93.3 |
| Llama 3.1 405B | 89.2 |
| Falcon 180B | 89.0 |
| Claude 3 Sonnet | 89.0 |
| DeepSeek-V3 | 88.9 |
| Gemini 1.0 Ultra | 87.8 |
| DeepSeek-V2 | 87.1 |
| PaLM 2-L | 86.8 |
| Gemini 1.5 Flash | 86.5 |
| Claude 3 Haiku | 85.9 |
| Llama 2 70B | 85.3 |
| Qwen2.5 72B | 84.8 |
| Gemini 1.0 Pro | 84.7 |
| LLaMA 65B | 84.2 |
| Stable Beluga 2 | 84.1 |
| PaLM 2-M | 84.0 |
| LLaMA 33B | 82.8 |
| Nemotron-4 15B | 82.4 |
| Phi-3 Medium 128K Instruct | 82.4 |
| PaLM 2-S | 82.0 |
| Mistral 7B v0.1 Base | 81.0 |
| Llama 2 13B | 80.7 |
| Qwen2.5-Coder-14B | 80.2 |
| Gopher (280B) | 79.2 |
| LLaMA-13B | 79.2 |
| OPT-175B | 79.1 |
| InternLM 20B | 78.1 |
| GPT-3 Davinci | 77.5 |
| Phi-3-Small-8K-Instruct | 77.0 |
| Qwen2.5 Coder 7B | 76.8 |
| Phi-3-mini-4k-instruct | 76.7 |
| Yi-9B | 76.4 |
| Yi-6B | 74.4 |
| XGen-7B-8K-Base | 74.2 |
| OpenLLaMA 7B | 71.8 |
| Gemma 2B | 71.4 |
| Qwen2.5-Coder-3B | 70.9 |