InstructGPT 6B
OpenAI · 2022-01-27
InstructGPT 6B is OpenAI's 6B-scale PPO-ptx policy from the InstructGPT work announced in January 2022. It was initialized from GPT-3, fine-tuned with supervised learning on human demonstrations, and further trained with RLHF via PPO while mixing pretraining-distribution gradients to reduce performance regressions.
Benchmark scores
| Benchmark | Score |
|---|---|
| BoolQ | 62.0 |
| GSM8K | 0.6 |
| HellaSwag (Unspecified Scoring Protocol) | 67.6 |