Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Abstract

Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique thataligns language models closely with human-centric values. The initial phase ofRLHF involves learning human values using a reward model from ranking data. Itis observed that the performance of the reward model degrades after one epochof training, and optimizing too much against the learned reward modeleventually hinders the true objective. This paper delves into these issues,leveraging the theoretical insights to design improved reward learningalgorithm termed 'Iterative Data Smoothing' (IDS). The core idea is that duringeach training epoch, we not only update the model with the data, but alsoupdate the date using the model, replacing hard labels with soft labels. Ourempirical findings highlight the superior performance of this approach over thetraditional methods.

Quick Read (beta)

loading the full paper ...