AlpacaEval 2.0 (length-controlled win rate)
Chat & Writing · 2024-04-01
AlpacaEval 2.0 measures instruction following by having GPT-4-Turbo compare a model's answers to a reference model's over 805 instructions, reported as a win rate. This row is the length-controlled win rate, the leaderboard's headline metric: a regression adjustment that removes the judge's preference for longer answers. The adjustment exists because the raw win rate is gameable -- a constant-output NullModel reaches 76.9 percent on it -- so the controlled figure is carried as the primary and the raw rate as a component. Being regression-adjusted, it lands on no item lattice and declares no denominator. A win rate against a reference has no guessing floor: an inferior model tends to zero, not to a chance level, so no chance_score is declared.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-4 Turbo | 55.0 |
| GPT-4 Turbo (1106 Preview) | 50.0 |
| Opus 3 | 40.5 |
| GPT-4 | 38.1 |
| Qwen2 72B Instruct | 38.1 |
| Qwen1.5 72B Chat | 36.6 |
| Claude 3 Sonnet | 34.9 |
| Llama 3 70B Instruct | 34.4 |
| Mistral Large 1.0 | 32.7 |
| Mixtral 8x22B Instruct | 30.9 |
| GPT-4 0613 | 30.2 |
| Claude 2 | 28.2 |
| Yi 34B Chat | 27.2 |
| DBRX Instruct | 25.4 |
| Claude 2.1 | 25.3 |
| Gemini 1.0 Pro | 24.4 |
| Qwen1.5 14B Chat | 23.9 |
| Mixtral 8x7B Instruct | 23.7 |
| Llama 3 8B Instruct | 22.9 |
| GPT-3.5 Turbo 16K 0613 | 22.7 |
| Tülu 2 DPO 70B | 21.2 |
| Mistral 7B Instruct v0.3 | 20.6 |
| GPT-3.5 Turbo 1106 | 19.3 |
| GPT-3.5 Turbo 0301 | 18.1 |
| Vicuna 33B v1.3 | 17.6 |
| Mistral 7B Instruct v0.2 | 17.1 |
| Qwen1.5 7B Chat | 14.7 |
| Llama 2 70B Chat | 14.7 |
| Vicuna 13B v1.3 | 10.8 |
| Gemma 7B IT | 10.4 |
| Llama 2 13B Chat | 8.4 |
| Falcon 40B Instruct | 5.6 |
| Gemma 2B IT | 5.4 |
| Llama 2 7B Chat | 5.4 |
| Qwen1.5 1.8B Chat | 2.6 |