advantage function
- Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models
- Risk-aware Direct Preference Optimization under Nested Risk Measure