Atlas

Benchmarks

← All benchmarks

Korean

General QA · 2024-06-01

Scale SEAL's Korean evaluation uses 1,000 prompts written for Korean speakers, with culturally fluent expert annotators assessing each response three times in parallel. Grading combines point-wise rubric scores with pairwise preferences on a seven-point Likert scale, aggregated into an Elo-style rating with confidence intervals.

Top models (higher is better)

ModelScore
o1 Preview66.4
Claude 3.7 Sonnet64.9
GPT-4o64.6
GPT-4.563.8
GPT-4 0125 Preview60.8
Gemini 1.5 Pro Experimental 082760.3
GPT-4o (2024-08-06)59.9
Claude 3.5 Sonnet (June 2024)59.4
Gemini 2.0 Flash 00155.6
Claude 3 Sonnet54.2
Opus 352.8
GPT-4o Mini51.7
GPT-4 061351.4
Llama 3.1 405B Instruct50.4
Mistral Large 2 (Instruct 2407)50.4
Gemini 1.5 Pro Preview 051440.4
Llama 3.1 70B Instruct37.2
C4AI Command R+ 08-202430.2
Llama 3.1 8B Instruct17.4
Loading Atlas data…