Atlas

Models

← All models

Llama 3.1 Nemotron 70B Instruct HF

NVIDIA · 2024-10-12 · 70.6B parameters

Llama 3.1 Nemotron 70B Instruct HF is NVIDIA's Hugging Face Transformers conversion of its helpfulness-aligned Llama 3.1 Nemotron 70B Instruct checkpoint, derived from Meta's Llama 3.1 70B Instruct. NVIDIA trained the source policy with REINFORCE RLHF using Llama-3.1-Nemotron-70B-Reward and HelpSteer2-Preference prompts; as of October 1, 2024, NVIDIA reported the source NeMo checkpoint first on Arena Hard, AlpacaEval 2 LC, and GPT-4-Turbo MT-Bench, while warning that this converted HF checkpoint can produce slightly different evaluation results. The converted checkpoint supports text-to-text generation with a 131,072-token context.

Benchmark scores

BenchmarkScore
AA-LCR7.0
AA-Omniscience Index-41.0
AIME 202511.0
Artificial Analysis Intelligence Index7.6
Artificial Analysis Omniscience Accuracy16.4
Artificial Analysis Omniscience Hallucination Rate68.8
Artificial Analysis Openness Index44.4
BBH (Open LLM Leaderboard v2)63.2
BRIDGE Medical (chain-of-thought)24.1
BRIDGE Medical (few-shot)42.8
BRIDGE Medical (zero-shot)32.8
Capability121.3
CritPt0.0
GPQA (Open LLM Leaderboard v2)25.8
GPQA Diamond46.5
Humanity's Last Exam (Text-Only)4.6
IFBench (Artificial Analysis)30.7
IFEval (Open LLM Leaderboard v2)73.8
LiveCodeBench (Artificial Analysis)16.9
MATH Level 5 (Open LLM Leaderboard v2)42.7
MATH-500 (Artificial Analysis source)73.3
MEDIC (clinical summarization)84.7
MEDIC (closed-ended)65.6
MEDIC (open-ended Elo)1897
MMLU Pro69.0
MMLU-Pro (Open LLM Leaderboard v2)49.2
MuSR (Open LLM Leaderboard v2)43.3
SciCode (Artificial Analysis)23.3
SnakeBench Average Apples2.9
SnakeBench Best Apples8.0
Loading Atlas data…