contextual bandits
A variation of multi-armed bandit problems where the decision of which action to take is influenced by the context provided at each round, allowing for more informed choices.
- Diffusion Models Meet Contextual Bandits
- Exploration via Feature Perturbation in Contextual Bandits
- Feel-Good Thompson Sampling for Contextual Bandits: a Markov Chain Monte Carlo Showdown
- Greedy Algorithms for Structured Bandits: A Sharp Characterization of Asymptotic Success / Failure
- Learning to Generalize: An Information Perspective on Neural Processes
- Scalable Exploration via Ensemble++
- Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
- Strategyproof Reinforcement Learning from Human Feedback
- Test-Time Scaling of Diffusion Models via Noise Trajectory Search
- Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits