regret guarantees
In reinforcement learning and decision theory, regret guarantees provide bounds on the difference between the reward achieved by the learning algorithm and the optimal reward, guiding the evaluation of algorithm performance.
- Constrained Linear Thompson Sampling
- Efficient Preference-Based Reinforcement Learning: Randomized Exploration meets Experimental Design
- Efficient Spectral Control of Partially Observed Linear Dynamical Systems
- Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
- Exploration via Feature Perturbation in Contextual Bandits
- When Lower-Order Terms Dominate: Adaptive Expert Algorithms for Heavy-Tailed Losses