entropy
- A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
- Off-policy Reinforcement Learning with Model-based Exploration Augmentation
- Preference Optimization by Estimating the Ratio of the Data Distribution
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning