Atlas

Benchmarks

← All benchmarks

PIQA

Classic NLP · 2019-11-26

A physical commonsense benchmark where models choose the more feasible solution to everyday problems.

Top models (higher is better)

ModelScore
GPT-4o Mini88.7
Phi-3.5-MoE-instruct88.6
Gemini 1.5 Flash 00287.5
Llama 3.1 405B85.9
Falcon 180B84.9
Inflection-184.2
DeepSeek-V283.9
Gemma 2 9B83.7
Mixtral 8x7B83.6
Mistral NeMo Base 240783.5
Stable Beluga 283.3
LLaMA 65B82.8
Qwen2.5 72B82.6
Nemotron-4 15B82.4
PaLM 540B82.3
Mistral 7B Instruct v0.182.2
Mistral 7B Instruct v0.282.2
Llama 2 34B81.9
MPT-30B81.9
Chinchilla81.8
Gopher (280B)81.8
Gemma 7B81.2
Llama 3.1 8B Instruct81.2
Phi-3.5 Mini Instruct81.0
PaLM 62B80.5
InternLM 20B80.3
LLaMA-13B80.1
Qwen 14B79.9
Baichuan2-13B-Base78.1
InternLM 7B77.9
Qwen 7B77.9
Vicuna 13B v1.177.4
Gemma 2B77.3
RedPajama INCITE 7B Base76.9
GPT-NeoX-20B76.7
Baichuan-7B76.2
OpenLLaMA 7B76.0
XGen-7B-8K-Base75.5
Dolly 2.0 12B75.4
Cerebras-GPT 13B73.5
Loading Atlas data…