Adversarial NLI
Classic NLP · 2019-10-31
A benchmark of adversarially collected natural language inference examples that tests robust textual reasoning under distribution shift.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-3.5 Turbo 1106 | 58.1 |
| Phi-3-Small-8K-Instruct | 58.1 |
| Llama 3 8B Instruct | 57.3 |
| Phi-3 Medium 128K Instruct | 55.8 |
| Mixtral 8x7B | 55.2 |
| Phi-3-mini-4k-instruct | 52.8 |
| Gemma 7B | 48.7 |
| Mistral 7B v0.1 Base | 47.1 |
| Phi-2 | 42.5 |