LAMBADA
Classic NLP · 2016-06-20
A long-context language modeling benchmark where the final word of a passage must be predicted from broader discourse.
Top models (higher is better)
| Model | Score |
|---|---|
| Falcon 180B | 79.8 |
| Llama 2 70B | 78.9 |
| Inflection-1 | 78.5 |
| PaLM 540B | 77.9 |
| LLaMA 65B | 77.7 |
| Chinchilla | 77.4 |
| Falcon 40B | 77.3 |
| LLaMA 33B | 77.2 |
| Llama 2 13B | 76.5 |
| LLaMA-13B | 75.2 |
| Falcon 7B | 74.9 |
| Gopher (280B) | 74.5 |
| Baichuan2-13B-Base | 74.0 |
| Baichuan2-7B-Base | 73.3 |
| Llama 2 7B | 73.3 |
| LLaMA 7B | 73.3 |
| InternLM 20B | 71.8 |
| Stable Beluga 2 | 71.3 |
| Qwen 14B | 71.1 |
| MPT-7B | 70.0 |
| Qwen 7B | 67.9 |
| InternLM 7B | 67.0 |
| Qwen 1.8B | 58.4 |
| ChatGLM2 6B (chat) | 54.3 |