Atlas

Benchmarks

← All benchmarks

WinoGrande

General QA · 2019-07-24

WinoGrande is a large-scale adversarial Winograd-style commonsense reasoning benchmark with 44k pronoun-resolution problems.

Top models (higher is better)

ModelScore
Llama 3.1 405B89.2
Opus 388.5
GPT-487.5
GPT-4 32K (0314)87.5
Falcon 180B87.1
DeepSeek-V286.3
DeepSeek-V385.2
Llama 3 70B83.5
PaLM 2-L83.0
Qwen2.5 72B82.3
Llama 3.1 405B Instruct82.2
GPT-3.5 Turbo 061381.6
Phi-3 Medium 128K Instruct81.5
Phi-3-Small-8K-Instruct81.5
Llama 2 70B80.2
PaLM 2-M79.2
Falcon2-11B78.3
Nemotron-4 15B78.0
PaLM 2-S77.9
LLaMA 65B77.0
PaLM 62B77.0
Qwen2.5-Coder-14B76.8
Llama 2 34B76.7
LLaMA 33B76.0
Llama 3 8B75.7
Claude 3 Sonnet75.1
Chinchilla74.9
Claude 3 Haiku74.2
Phi-1.573.4
LLaMA-13B73.0
Yi-9B73.0
DeepSeek-Coder-V2-Lite-Base72.9
Qwen2.5 Coder 7B72.9
Llama 2 13B72.8
Yi-6B71.3
MPT-30B71.0
Phi-3-mini-4k-instruct70.8
Vicuna 13B v1.170.8
GPT-3.5 Turbo 110668.8
Qwen2.5-Coder-3B67.4
Loading Atlas data…