Atlas

Benchmarks

← All benchmarks

Global PIQA

General QA · 2025-10-28

Global PIQA is a hand-written physical commonsense reasoning benchmark covering 141 language varieties across 116 countries, built by more than 300 researchers and verified by native speakers. This definition holds source-reported aggregate accuracy when a model card does not identify a split or scoring protocol; exact split/protocol measurements use child benchmark slugs.

Top models (higher is better)

ModelScore
Gemini 3 Pro Preview93.4
Gemini 3 Flash Preview92.8
Gemini 2.5 Pro91.5
GPT-5.291.2
GPT-5.190.9
Gemini 2.5 Flash90.2
Sonnet 4.590.1
Grok 4.1 Fast85.6
Loading Atlas data…