Atlas

Benchmarks

← All benchmarks

Global PIQA - Parallel (Log-Likelihood Accuracy)

General QA · 2025-10-28

Global PIQA's parallel split presents aligned physical-commonsense questions with four answer options. This definition compares answer-option log likelihoods without requiring the model to generate a choice token.

Top models (higher is better)

ModelScore
Llama 3.1 70B30.7
Gemma 2 27B30.7
Qwen2.5 72B29.8
Gemma 3 12B PT29.3
Qwen2.5 32B27.7
Gemma 2 9B27.3
Mixtral 8x7B25.2
Llama 3.1 8B25.2
Falcon 40B24.5
Mistral 7B v0.1 Base23.5
Falcon 7B21.8
Gemma 3 270M21.1
Loading Atlas data…