reinforcement learning framework
A paradigm of machine learning where agents learn to make decisions by receiving rewards or punishments based on their actions, facilitating learning through trial and error.
- AutoEdit: Automatic Hyperparameter Tuning for Image Editing
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
- R$^2$ec: Towards Large Recommender Models with Reasoning
- Reward Reasoning Models
- Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information Foraging
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning