AlpacaEval 2.0 (raw win rate)
Chat & Writing · 2024-04-01
The raw, uncontrolled AlpacaEval 2.0 win rate against the reference model, judged by GPT-4-Turbo over 805 instructions. This is the weighted rate, averaging the judge's per-comparison preference probability rather than counting discrete wins, so it is a mean of per-item scores over the 805 comparisons rather than a pass count. Carried as a component of the length-controlled factor group: it re-measures the same comparisons without the length adjustment, and scores several points lower for verbose models as a result. A handful of models were judged on 801 to 804 comparisons instead of the full 805 because individual generations failed; the declared count is the full set and each row records its own total in the notes.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-4 Turbo (1106 Preview) | 50.0 |
| GPT-4 Turbo | 46.1 |
| Llama 3 70B Instruct | 33.2 |
| Qwen2 72B Instruct | 29.9 |
| Yi 34B Chat | 29.7 |
| Opus 3 | 29.1 |
| Qwen1.5 72B Chat | 26.5 |
| Claude 3 Sonnet | 25.6 |
| GPT-4 | 23.6 |
| Llama 3 8B Instruct | 22.6 |
| Mixtral 8x22B Instruct | 22.2 |
| Mistral Large 1.0 | 21.4 |
| Qwen1.5 14B Chat | 18.6 |
| DBRX Instruct | 18.4 |
| Mixtral 8x7B Instruct | 18.3 |
| Gemini 1.0 Pro | 18.2 |
| Claude 2 | 17.2 |
| Mistral 7B Instruct v0.3 | 16.7 |
| Tülu 2 DPO 70B | 16.0 |
| GPT-4 0613 | 15.8 |
| Claude 2.1 | 15.7 |
| Mistral 7B Instruct v0.2 | 14.7 |
| GPT-3.5 Turbo 16K 0613 | 14.1 |
| Llama 2 70B Chat | 13.9 |
| Vicuna 33B v1.3 | 12.7 |
| Qwen1.5 7B Chat | 11.8 |
| GPT-3.5 Turbo 0301 | 9.6 |
| GPT-3.5 Turbo 1106 | 9.2 |
| Llama 2 13B Chat | 7.7 |
| Vicuna 13B v1.3 | 7.1 |
| Gemma 7B IT | 6.9 |
| Llama 2 7B Chat | 5.0 |
| Qwen1.5 1.8B Chat | 3.7 |
| Gemma 2B IT | 3.4 |
| Falcon 40B Instruct | 3.3 |