CommonSenseQA 2
Classic NLP · 2022-01-14
A harder, bias-reduced multiple-choice benchmark that probes everyday commonsense beyond lexical shortcuts.
Top models (higher is better)
| Model | Score |
|---|---|
| T5 11B | 67.8 |
| T5 3B | 60.2 |
| GPT-3.5 Turbo 0613 | 57.0 |
| Llama 2 70B | 50.0 |