verifiable rewards
Rewards in a reinforcement learning framework that can be easily verified against a ground truth or expected outcomes. This is crucial for ensuring the reliability and consistency of learned policies.
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Generalizing Verifiable Instruction Following
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- RLVR-World: Training World Models with Reinforcement Learning
- Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
- Rethinking Verification for LLM Code Generation: From Generation to Testing
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards