PIQA
Classic NLP · 2019-11-26
A physical commonsense benchmark where models choose the more feasible solution to everyday problems.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-4o Mini | 88.7 |
| Phi-3.5-MoE-instruct | 88.6 |
| Gemini 1.5 Flash 002 | 87.5 |
| Llama 3.1 405B | 85.9 |
| Falcon 180B | 84.9 |
| Inflection-1 | 84.2 |
| DeepSeek-V2 | 83.9 |
| Gemma 2 9B | 83.7 |
| Mixtral 8x7B | 83.6 |
| Mistral NeMo Base 2407 | 83.5 |
| Stable Beluga 2 | 83.3 |
| LLaMA 65B | 82.8 |
| Qwen2.5 72B | 82.6 |
| Nemotron-4 15B | 82.4 |
| PaLM 540B | 82.3 |
| Mistral 7B Instruct v0.1 | 82.2 |
| Mistral 7B Instruct v0.2 | 82.2 |
| Llama 2 34B | 81.9 |
| MPT-30B | 81.9 |
| Chinchilla | 81.8 |
| Gopher (280B) | 81.8 |
| Gemma 7B | 81.2 |
| Llama 3.1 8B Instruct | 81.2 |
| Phi-3.5 Mini Instruct | 81.0 |
| PaLM 62B | 80.5 |
| InternLM 20B | 80.3 |
| LLaMA-13B | 80.1 |
| Qwen 14B | 79.9 |
| Baichuan2-13B-Base | 78.1 |
| InternLM 7B | 77.9 |
| Qwen 7B | 77.9 |
| Vicuna 13B v1.1 | 77.4 |
| Gemma 2B | 77.3 |
| RedPajama INCITE 7B Base | 76.9 |
| GPT-NeoX-20B | 76.7 |
| Baichuan-7B | 76.2 |
| OpenLLaMA 7B | 76.0 |
| XGen-7B-8K-Base | 75.5 |
| Dolly 2.0 12B | 75.4 |
| Cerebras-GPT 13B | 73.5 |