process reward models
Models that help predict or estimate the rewards associated with different actions in a given process, often used in reinforcement learning setups to inform decision-making strategies effectively.
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- DreamPRM: Domain-reweighted Process Reward Model for Multimodal Reasoning
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
- Know What You Don't Know: Uncertainty Calibration of Process Reward Models
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
- Reasoning Is Not a Race: When Stopping Early Beats Going Deeper
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
- Unlocking Multimodal Mathematical Reasoning via Process Reward Model
- Value-Guided Search for Efficient Chain-of-Thought Reasoning