Atlas

Benchmarks

← All benchmarks

Arabic

General QA · 2024-12-18

Elo rating on Scale SEAL's Arabic multilingual prompt set.

Top models (higher is better)

ModelScore
Gemini 1.5 Pro Experimental 08271147
Gemini Exp-12061138
Gemini 2.0 Flash Thinking Experimental 01-211120
O11120
Gemini Experimental 11211116
o3-mini1093
Gemini 2.0 Flash Experimental1090
o1 Preview1087
GPT-4o (2024-08-06)1066
Aya Expanse 32B1025
GPT-4 Turbo (1106 Preview)1011
Claude 3.5 Sonnet (June 2024)995
Mistral Large 2 (Instruct 2407)970
Gemini 1.5 Flash Preview 0514967
Aya-23-35B932
Llama 3.1 405B Instruct875
Llama 3.3 70B Instruct808
Gemma 2 27B661
Loading Atlas data…