Atlas

Benchmarks

← All benchmarks

TriviaQA

Classic NLP · 2017-05-09

An open-domain question answering benchmark with challenging trivia questions paired with evidence documents.

Top models (higher is better)

ModelScore
Claude 287.5
Claude 1.386.7
PaLM 2-L86.1
GPT-3.5 Turbo 110685.8
GPT-4 061384.8
DeepSeek-V382.9
Llama 3.1 405B82.7
Mixtral 8x7B82.2
PaLM 2-M81.7
DeepSeek-V280.0
Claude Instant78.9
Claude Instant 1.278.7
PaLM 2-S75.2
Phi-3 Medium 128K Instruct73.9
Qwen2.5 72B71.9
Llama 3 8B Instruct67.7
Phi-3-mini-4k-instruct64.0
Phi-3-Small-8K-Instruct58.1
Gemma 2B53.2
Phi-245.2
Loading Atlas data…