PaLM 2-L
Google DeepMind · 2023-05-10
PaLM 2-L, also referred to by Google as PaLM 2-L-Unicorn, is the largest variant evaluated in Google's PaLM 2 technical report. It is a Transformer-based text language model trained with a tuned mixture of pre-training objectives and a more multilingual, diverse corpus than PaLM; Google reports that it is significantly smaller than PaLM-540B yet uses more training compute. Google withholds exact model-size and architecture details, and the report averages L results across its last five checkpoints rather than naming a single public checkpoint.
Benchmark scores
| Benchmark | Score |
|---|---|
| ARC AI2 | 69.2 |
| BBH (BIG-Bench Hard) | 77.7 |
| BoolQ | 90.9 |
| Capability | 116.0 |
| Epoch Capabilities Index (ECI) | 114 |
| GSM8K | 80.0 |
| HellaSwag (Unspecified Scoring Protocol) | 86.8 |
| MATH | 34.4 |
| MGSM | 74.7 |
| MMLU | 78.4 |
| TriviaQA | 86.1 |
| WinoGrande | 83.0 |