reward models
Models that estimate the value or reward associated with various states or actions in reinforcement learning, guiding agent behavior and decision-making.
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
- EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
- Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback
- HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
- Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning
- Incentivizing LLMs to Self-Verify Their Answers
- LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- Majority of the Bests: Improving Best-of-N via Bootstrapping
- Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability
- Policy Optimized Text-to-Image Pipeline Design
- Preference Optimization by Estimating the Ratio of the Data Distribution
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
- Reward Reasoning Models
- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods
- Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty