NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
reward feedback
3 papers
Constrained Feedback Learning for Non-Stationary Multi-Armed Bandits
Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs