Atlas

Benchmarks

← All benchmarks

ScienceQA

Multimodal · 2022-09-20

A multimodal multiple-choice science benchmark combining text, images, and diagrams with rich rationales.

Top models (higher is better)

ModelScore
Phi-3.5-vision-instruct91.3
GPT-4o88.5
Gemini 1.0 Pro Vision79.7
LLaVA v1.6 Mistral 7B72.8
Claude 3 Haiku72.0
Flan-T5 XXL67.4
LLaVA v1.5 7B66.8
Llama 2 13B55.8
LLaMA-13B43.3
Llama 2 7B43.1
LLaMA 7B36.2
Loading Atlas data…