human feedback
In the context of AI, human feedback refers to input given by human annotators or users that helps to fine-tune model outputs or guide the training process. It is often used in reinforcement learning from human feedback (RLHF) frameworks to improve model alignment with human preferences and understanding.
- A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- Avoiding exp(R) scaling in RLHF through Preference-based Exploration
- Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
- Capturing Individual Human Preferences with Reward Features
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Doubly Robust Alignment for Large Language Models
- Efficient Preference-Based Reinforcement Learning: Randomized Exploration meets Experimental Design
- Efficient and Near-Optimal Algorithm for Contextual Dueling Bandits with Offline Regression Oracles
- Evaluating LLM-contaminated Crowdsourcing Data Without Ground Truth
- Explainable Reinforcement Learning from Human Feedback to Improve Alignment
- Greedy Sampling Is Provably Efficient For RLHF
- Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
- Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning
- Improving Video Generation with Human Feedback
- Inference-time Alignment in Continuous Space
- Information-Theoretic Reward Decomposition for Generalizable RLHF
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
- Progress Reward Model for Reinforcement Learning via Large Language Models
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization
- Scalable Valuation of Human Feedback through Provably Robust Model Alignment
- Stochastically Dominant Peer Prediction
- Strategyproof Reinforcement Learning from Human Feedback
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
- VPO: Reasoning Preferences Optimization Based on $\mathcal{V}$-Usable Information
- What Makes a Reward Model a Good Teacher? An Optimization Perspective