advantage estimation
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
- Real-World Reinforcement Learning of Active Perception Behaviors
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models