OpenBookQA
Science · 2018-09-08
OpenBookQA is a multiple-choice elementary science question-answering benchmark requiring use of a small open-book fact set plus common knowledge. It is often used to measure scientific commonsense reasoning.
Top models (higher is better)
| Model | Score |
|---|---|
| Phi-3-Small-8K-Instruct | 88.0 |
| Phi-3 Medium 128K Instruct | 87.4 |
| GPT-3.5 Turbo 1106 | 86.0 |
| Mixtral 8x7B | 85.8 |
| Mistral 7B v0.1 Base | 79.8 |
| Gemma 7B | 78.6 |
| Phi-2 | 73.6 |
| Falcon 180B | 64.2 |
| Llama 2 70B | 60.2 |
| LLaMA 65B | 60.2 |
| LLaMA 33B | 58.6 |
| Llama 2 34B | 58.2 |
| PaLM 2-M | 57.4 |
| Llama 2 13B | 57.0 |
| Falcon 40B | 56.6 |
| LLaMA-13B | 56.4 |
| MPT-30B | 52.0 |
| PaLM 62B | 50.4 |
| XGen-7B-8K-Base | 40.2 |
| RedPajama INCITE 7B Base | 40.0 |
| Dolly 2.0 12B | 39.2 |
| OpenLLaMA 7B | 39.0 |
| GPT-NeoX-20B | 38.8 |
| Phi-1.5 | 37.2 |
| Cerebras-GPT 13B | 35.8 |
| Vicuna 13B v1.1 | 33.0 |
| StableLM Tuned Alpha 7B | 32.4 |
| OPT-1.3B | 24.0 |
| GPT-2 (1.5B) | 22.4 |