Atlas

Benchmarks

← All benchmarks

ARC AI2

Classic NLP · 2018-03-14

A multiple-choice science benchmark assessing grade-school science knowledge and reasoning on questions from the AI2 ARC dataset.

Top models (higher is better)

ModelScore
GPT-496.3
DeepSeek-V395.3
Llama 3.1 405B95.3
Qwen2.5 72B94.5
DeepSeek-V292.2
Phi-3 Medium 128K Instruct91.6
Phi-3-Small-8K-Instruct90.7
GPT-3.5 Turbo 110687.4
Claude Instant 1.286.3
Stable Beluga 286.1
Claude Instant85.7
Phi-3-mini-4k-instruct84.9
Qwen 14B84.4
Llama 3 8B Instruct82.8
InternLM 20B81.7
Phi-275.9
Qwen 7B75.3
InternLM 7B69.5
PaLM 2-L69.2
Qwen2.5-Coder-14B66.0
PaLM 2-M64.9
ChatGLM2 6B (chat)61.0
Qwen2.5 Coder 7B60.9
PaLM 2-S59.6
DeepSeek-Coder-V2-Lite-Base57.3
Yi-9B55.6
Nemotron-4 15B55.5
Llama 2 34B54.5
Qwen 1.8B53.2
Qwen2.5-Coder-3B52.9
LLaMA-13B52.7
PaLM 62B52.5
MPT-30B50.6
Yi-6B50.3
StarCoder2-15B47.2
Qwen2.5 Coder 1.5B45.2
Phi-1.544.4
Vicuna 13B v1.143.2
DeepSeek-Coder-Base 33B42.2
Gemma 2B42.1
Loading Atlas data…