Atlas

Benchmarks

← All benchmarks

Global PIQA - Non-Parallel (Strict Exact Match)

General QA · 2025-10-28

Global PIQA's non-parallel split contains locally authored physical-commonsense questions with two answer options. This definition scores generated answers using exact match with strict answer extraction.

Top models (higher is better)

ModelScore
Gemma 4 31B IT84.9
Gemma 3 27B IT81.7
Command A81.2
Qwen3.5-27B78.5
Qwen2.5 72B Instruct78.3
Gemma 3 12B IT78.1
gpt-oss-20b77.8
Qwen3 14B77.1
Llama 3.1 70B Instruct76.2
Qwen3 32B75.9
Qwen2.5 Instruct 32B75.2
Apertus 70B Instruct74.3
Phi-474.2
Qwen3 8B74.0
Qwen3.5-9B73.5
Gemma 2 27B IT73.4
Qwen2.5 14B Instruct73.4
DeepSeek-R1-Distill-Qwen-32B72.2
Aya Expanse 32B70.9
DeepSeek R1 Distill Qwen 14B70.7
Qwen3 4B70.6
gpt-oss-120b69.5
C4AI Command R 08-202468.9
Gemma 2 9B IT68.3
Gemma 3 4B IT67.6
Apertus 8B Instruct67.0
Qwen2.5 7B Instruct66.5
C4AI Command R+ 08-202465.8
Llama 3.1 8B Instruct61.0
Qwen3 1.7B60.8
Qwen2.5 3B Instruct60.2
Phi-3 Medium 4K Instruct59.5
Tiny Aya Global59.5
Command R7B (Dec 2024)59.0
Phi-4-mini-instruct58.5
Gemma 2 2B IT58.4
Phi-3.5 Mini Instruct56.9
Phi-3-mini-4k-instruct55.2
Gemma 3 1B IT54.6
Mistral 7B Instruct v0.354.6
Loading Atlas data…