WinoGrande
General QA · 2019-07-24
WinoGrande is a large-scale adversarial Winograd-style commonsense reasoning benchmark with 44k pronoun-resolution problems.
Top models (higher is better)
| Model | Score |
|---|---|
| Llama 3.1 405B | 89.2 |
| Opus 3 | 88.5 |
| GPT-4 | 87.5 |
| GPT-4 32K (0314) | 87.5 |
| Falcon 180B | 87.1 |
| DeepSeek-V2 | 86.3 |
| DeepSeek-V3 | 85.2 |
| Llama 3 70B | 83.5 |
| PaLM 2-L | 83.0 |
| Qwen2.5 72B | 82.3 |
| Llama 3.1 405B Instruct | 82.2 |
| GPT-3.5 Turbo 0613 | 81.6 |
| Phi-3 Medium 128K Instruct | 81.5 |
| Phi-3-Small-8K-Instruct | 81.5 |
| Llama 2 70B | 80.2 |
| PaLM 2-M | 79.2 |
| Falcon2-11B | 78.3 |
| Nemotron-4 15B | 78.0 |
| PaLM 2-S | 77.9 |
| LLaMA 65B | 77.0 |
| PaLM 62B | 77.0 |
| Qwen2.5-Coder-14B | 76.8 |
| Llama 2 34B | 76.7 |
| LLaMA 33B | 76.0 |
| Llama 3 8B | 75.7 |
| Claude 3 Sonnet | 75.1 |
| Chinchilla | 74.9 |
| Claude 3 Haiku | 74.2 |
| Phi-1.5 | 73.4 |
| LLaMA-13B | 73.0 |
| Yi-9B | 73.0 |
| DeepSeek-Coder-V2-Lite-Base | 72.9 |
| Qwen2.5 Coder 7B | 72.9 |
| Llama 2 13B | 72.8 |
| Yi-6B | 71.3 |
| MPT-30B | 71.0 |
| Phi-3-mini-4k-instruct | 70.8 |
| Vicuna 13B v1.1 | 70.8 |
| GPT-3.5 Turbo 1106 | 68.8 |
| Qwen2.5-Coder-3B | 67.4 |