preference-based reinforcement learning
Preference-based reinforcement learning leverages user feedback or preferences instead of explicit reward signals to guide the learning process, allowing agents to adapt their behavior based on subjective evaluations of outcomes.
- Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning
- MisoDICE: Multi-Agent Imitation from Mixed-Quality Demonstrations
- PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models
- PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models
- STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
- STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization
- Trajectory Graph Learning: Aligning with Long Trajectories in Reinforcement Learning Without Reward Design
- Uncertainty-aware Preference Alignment for Diffusion Policies