Atlas

Benchmarks

← All benchmarks

LAMBADA

Classic NLP · 2016-06-20

A long-context language modeling benchmark where the final word of a passage must be predicted from broader discourse.

Top models (higher is better)

ModelScore
Falcon 180B79.8
Llama 2 70B78.9
Inflection-178.5
PaLM 540B77.9
LLaMA 65B77.7
Chinchilla77.4
Falcon 40B77.3
LLaMA 33B77.2
Llama 2 13B76.5
LLaMA-13B75.2
Falcon 7B74.9
Gopher (280B)74.5
Baichuan2-13B-Base74.0
Baichuan2-7B-Base73.3
Llama 2 7B73.3
LLaMA 7B73.3
InternLM 20B71.8
Stable Beluga 271.3
Qwen 14B71.1
MPT-7B70.0
Qwen 7B67.9
InternLM 7B67.0
Qwen 1.8B58.4
ChatGLM2 6B (chat)54.3
Loading Atlas data…