Nemotron-4 15B
NVIDIA · 2024-02-26
Nemotron-4 15B is NVIDIA's 15B-class multilingual decoder-only language model, introduced in a technical report on February 26, 2024. It was pretrained on 8 trillion tokens—70% English, 15% multilingual text across 53 additional natural languages, and 15% source code across 43 programming languages—and then underwent continued training on an undisclosed small number of additional tokens under shifted data distributions, so its exact cumulative training exposure is unknown. Its 32-layer architecture uses a 4,096-token sequence length, GQA, RoPE, LayerNorm, squared ReLU, and a 256,000-token SentencePiece BPE vocabulary; NVIDIA reported leading multilingual performance at its scale.
Benchmark scores
| Benchmark | Score |
|---|---|
| ARC AI2 | 55.5 |
| BBH (BIG-Bench Hard) | 58.7 |
| Capability | 103.0 |
| Epoch Capabilities Index (ECI) | 107 |
| GSM8K | 46.0 |
| HellaSwag (Unspecified Scoring Protocol) | 82.4 |
| MMLU | 58.7 |
| PIQA | 82.4 |
| WinoGrande | 78.0 |