reinforcement learning from human feedback

Reinforcement learning from human feedback (RLHF) is an approach where human insights or evaluations are used to guide and improve the learning of an AI model. This method helps models align their outcomes with human preferences, often enhancing interpretability and usefulness.

12 papers