Smaug-72B-v0.1
Abacus.AI · 2024-02-02 · 72.3B parameters
Smaug-72B-v0.1 is Abacus.AI's 72B instruction-tuned model fine-tuned from the Qwen-72B-derived moreh/MoMo-72B-lora-1.8.7-DPO checkpoint using DPO-Positive (DPOP). DPOP addresses a standard-DPO failure mode in which preferred-completion likelihood can decrease on response pairs with low edit distance; at release, Smaug-72B-v0.1 ranked first on the Open LLM Leaderboard.