Atlas

Benchmarks

← All benchmarks

Global PIQA - Parallel (Strict Exact Match)

General QA · 2025-10-28

Global PIQA's parallel split presents aligned physical-commonsense questions with four answer options. This definition scores generated answers using exact match with strict answer extraction.

Top models (higher is better)

ModelScore
Gemma 4 31B IT82.4
Qwen3.5-27B71.6
gpt-oss-20b65.6
Gemma 3 27B IT65.4
Command A65.0
Qwen3 14B64.1
Qwen2.5 72B Instruct60.8
Qwen3 32B60.8
Gemma 3 12B IT59.2
Qwen3.5-9B58.7
Qwen3 8B58.4
DeepSeek-R1-Distill-Qwen-32B56.5
Qwen2.5 Instruct 32B56.4
Gemma 2 27B IT54.5
Qwen3 4B53.0
DeepSeek R1 Distill Qwen 14B51.8
Llama 3.1 70B Instruct51.1
Qwen2.5 14B Instruct49.0
gpt-oss-120b48.0
Gemma 2 9B IT47.4
Apertus 70B Instruct47.1
Phi-446.5
Aya Expanse 32B45.9
Gemma 3 4B IT42.8
Qwen2.5 7B Instruct42.1
Qwen3 1.7B41.5
C4AI Command R 08-202440.7
Apertus 8B Instruct38.9
C4AI Command R+ 08-202438.1
Llama 3.1 8B Instruct35.8
Phi-3 Medium 4K Instruct35.6
Qwen2.5 3B Instruct35.4
Gemma 2 2B IT34.1
Tiny Aya Global33.7
Phi-4-mini-instruct33.5
Command R7B (Dec 2024)33.0
Phi-3.5 Mini Instruct32.0
Qwen3 0.6B30.3
Phi-3-mini-4k-instruct29.0
Llama 3.2 3B Instruct27.4
Loading Atlas data…