InstructGPT 1.3B
OpenAI · 2022-01-27
InstructGPT 1.3B is a 1.3-billion-parameter instruction-following model from OpenAI's January 2022 InstructGPT work. It was trained using supervised fine-tuning on human demonstrations and reinforcement learning from human feedback (RLHF), and demonstrated that smaller RLHF-tuned models can outperform much larger untuned GPT-3 models on instruction-following tasks.
Benchmark scores
| Benchmark | Score |
|---|---|
| BoolQ | 45.1 |
| GSM8K | 0.0 |
| HellaSwag (Unspecified Scoring Protocol) | 56.1 |