rlhf
Reinforcement Learning from Human Feedback (RLHF) is an approach that integrates human feedback into the reinforcement learning process to guide and improve learning and decision-making in AI agents.
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning
- Beyond Expectations: Quantile-Guided Alignment for Risk-Calibrated Language Models
- Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
- LLM Safety Alignment is Divergence Estimation in Disguise