Tülu 2 DPO 70B
Allen Institute for AI (Ai2) · 2023-11-17 · 69B parameters
Tülu 2 DPO 70B is Ai2's instruction-following Llama 2 70B checkpoint, first supervised-fine-tuned on the Tülu V2 mixture and then further aligned with Direct Preference Optimization on a filtered, binarized version of UltraFeedback. Ai2 publicly released the checkpoint, data, and training and evaluation code in November 2023.