RLHF at scale; the recipe behind instruction-following assistants.
Read the paper ↗
Reinforcement learning from human feedback makes a smaller model preferred over a far larger raw one.