Atlas

Benchmarks

← All benchmarks

CommonSenseQA 2

Classic NLP · 2022-01-14

A harder, bias-reduced multiple-choice benchmark that probes everyday commonsense beyond lexical shortcuts.

Top models (higher is better)

ModelScore
T5 11B67.8
T5 3B60.2
GPT-3.5 Turbo 061357.0
Llama 2 70B50.0
Loading Atlas data…