Atlas

Benchmarks

← All benchmarks

AlpacaEval 2.0 (raw win rate)

Chat & Writing · 2024-04-01

The raw, uncontrolled AlpacaEval 2.0 win rate against the reference model, judged by GPT-4-Turbo over 805 instructions. This is the weighted rate, averaging the judge's per-comparison preference probability rather than counting discrete wins, so it is a mean of per-item scores over the 805 comparisons rather than a pass count. Carried as a component of the length-controlled factor group: it re-measures the same comparisons without the length adjustment, and scores several points lower for verbose models as a result. A handful of models were judged on 801 to 804 comparisons instead of the full 805 because individual generations failed; the declared count is the full set and each row records its own total in the notes.

Top models (higher is better)

ModelScore
GPT-4 Turbo (1106 Preview)50.0
GPT-4 Turbo46.1
Llama 3 70B Instruct33.2
Qwen2 72B Instruct29.9
Yi 34B Chat29.7
Opus 329.1
Qwen1.5 72B Chat26.5
Claude 3 Sonnet25.6
GPT-423.6
Llama 3 8B Instruct22.6
Mixtral 8x22B Instruct22.2
Mistral Large 1.021.4
Qwen1.5 14B Chat18.6
DBRX Instruct18.4
Mixtral 8x7B Instruct18.3
Gemini 1.0 Pro18.2
Claude 217.2
Mistral 7B Instruct v0.316.7
Tülu 2 DPO 70B16.0
GPT-4 061315.8
Claude 2.115.7
Mistral 7B Instruct v0.214.7
GPT-3.5 Turbo 16K 061314.1
Llama 2 70B Chat13.9
Vicuna 33B v1.312.7
Qwen1.5 7B Chat11.8
GPT-3.5 Turbo 03019.6
GPT-3.5 Turbo 11069.2
Llama 2 13B Chat7.7
Vicuna 13B v1.37.1
Gemma 7B IT6.9
Llama 2 7B Chat5.0
Qwen1.5 1.8B Chat3.7
Gemma 2B IT3.4
Falcon 40B Instruct3.3
Loading Atlas data…