Atlas

Benchmarks

← All benchmarks

HellaSwag (Unspecified Scoring Protocol)

General QA · 2019-05-19

HellaSwag results for which the reporting source does not establish whether raw or length-normalized multiple-choice scoring was used. Values are retained as reported but quarantined from protocol-specific comparisons.

Top models (higher is better)

ModelScore
Opus 395.4
GPT-495.3
Gemini 1.5 Pro93.3
Claude 3 Sonnet89.0
Gemini 1.0 Ultra87.8
PaLM 2-L86.8
Gemini 1.5 Flash86.5
Claude 3 Haiku85.9
Gemini 1.0 Pro84.7
Loading Atlas data…