Atlas

Benchmarks

← All benchmarks

Adversarial NLI

Classic NLP · 2019-10-31

A benchmark of adversarially collected natural language inference examples that tests robust textual reasoning under distribution shift.

Top models (higher is better)

ModelScore
GPT-3.5 Turbo 110658.1
Phi-3-Small-8K-Instruct58.1
Llama 3 8B Instruct57.3
Phi-3 Medium 128K Instruct55.8
Mixtral 8x7B55.2
Phi-3-mini-4k-instruct52.8
Gemma 7B48.7
Mistral 7B v0.1 Base47.1
Phi-242.5
Loading Atlas data…