Atlas

Benchmarks

← All benchmarks

OpenBookQA

Science · 2018-09-08

OpenBookQA is a multiple-choice elementary science question-answering benchmark requiring use of a small open-book fact set plus common knowledge. It is often used to measure scientific commonsense reasoning.

Top models (higher is better)

ModelScore
Phi-3-Small-8K-Instruct88.0
Phi-3 Medium 128K Instruct87.4
GPT-3.5 Turbo 110686.0
Mixtral 8x7B85.8
Mistral 7B v0.1 Base79.8
Gemma 7B78.6
Phi-273.6
Falcon 180B64.2
Llama 2 70B60.2
LLaMA 65B60.2
LLaMA 33B58.6
Llama 2 34B58.2
PaLM 2-M57.4
Llama 2 13B57.0
Falcon 40B56.6
LLaMA-13B56.4
MPT-30B52.0
PaLM 62B50.4
XGen-7B-8K-Base40.2
RedPajama INCITE 7B Base40.0
Dolly 2.0 12B39.2
OpenLLaMA 7B39.0
GPT-NeoX-20B38.8
Phi-1.537.2
Cerebras-GPT 13B35.8
Vicuna 13B v1.133.0
StableLM Tuned Alpha 7B32.4
OPT-1.3B24.0
GPT-2 (1.5B)22.4
Loading Atlas data…