Llama 3.1 Nemotron 70B Instruct HF
NVIDIA · 2024-10-12 · 70.6B parameters
Llama 3.1 Nemotron 70B Instruct HF is NVIDIA's Hugging Face Transformers conversion of its helpfulness-aligned Llama 3.1 Nemotron 70B Instruct checkpoint, derived from Meta's Llama 3.1 70B Instruct. NVIDIA trained the source policy with REINFORCE RLHF using Llama-3.1-Nemotron-70B-Reward and HelpSteer2-Preference prompts; as of October 1, 2024, NVIDIA reported the source NeMo checkpoint first on Arena Hard, AlpacaEval 2 LC, and GPT-4-Turbo MT-Bench, while warning that this converted HF checkpoint can produce slightly different evaluation results. The converted checkpoint supports text-to-text generation with a 131,072-token context.