multi-armed bandit
A problem formulation in decision theory and reinforcement learning where an agent must choose between multiple options with uncertain rewards, aiming to maximize total gain.
- Balancing Performance and Costs in Best Arm Identification
- Bandit Guided Submodular Curriculum for Adaptive Subset Selection
- Distributed Multi-Agent Bandits Over Erdős-Rényi Random Networks
- Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
- LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits
- Learning Across the Gap: Hybrid Multi-armed Bandits with Heterogeneous Offline and Online Data
- Optimal Estimation of the Best Mean in Multi-Armed Bandits
- Oracle-Efficient Combinatorial Semi-Bandits
- Precise Asymptotics and Refined Regret of Variance-Aware UCB