Atlas

Benchmarks

← All benchmarks

AlpacaEval 2.0 (length-controlled win rate)

Chat & Writing · 2024-04-01

AlpacaEval 2.0 measures instruction following by having GPT-4-Turbo compare a model's answers to a reference model's over 805 instructions, reported as a win rate. This row is the length-controlled win rate, the leaderboard's headline metric: a regression adjustment that removes the judge's preference for longer answers. The adjustment exists because the raw win rate is gameable -- a constant-output NullModel reaches 76.9 percent on it -- so the controlled figure is carried as the primary and the raw rate as a component. Being regression-adjusted, it lands on no item lattice and declares no denominator. A win rate against a reference has no guessing floor: an inferior model tends to zero, not to a chance level, so no chance_score is declared.

Top models (higher is better)

ModelScore
GPT-4 Turbo55.0
GPT-4 Turbo (1106 Preview)50.0
Opus 340.5
GPT-438.1
Qwen2 72B Instruct38.1
Qwen1.5 72B Chat36.6
Claude 3 Sonnet34.9
Llama 3 70B Instruct34.4
Mistral Large 1.032.7
Mixtral 8x22B Instruct30.9
GPT-4 061330.2
Claude 228.2
Yi 34B Chat27.2
DBRX Instruct25.4
Claude 2.125.3
Gemini 1.0 Pro24.4
Qwen1.5 14B Chat23.9
Mixtral 8x7B Instruct23.7
Llama 3 8B Instruct22.9
GPT-3.5 Turbo 16K 061322.7
Tülu 2 DPO 70B21.2
Mistral 7B Instruct v0.320.6
GPT-3.5 Turbo 110619.3
GPT-3.5 Turbo 030118.1
Vicuna 33B v1.317.6
Mistral 7B Instruct v0.217.1
Qwen1.5 7B Chat14.7
Llama 2 70B Chat14.7
Vicuna 13B v1.310.8
Gemma 7B IT10.4
Llama 2 13B Chat8.4
Falcon 40B Instruct5.6
Gemma 2B IT5.4
Llama 2 7B Chat5.4
Qwen1.5 1.8B Chat2.6
Loading Atlas data…