upper confidence bound
A strategy in bandit problems where the algorithm calculates an upper confidence limit for the expected rewards of each action, balancing exploration and exploitation.
- Gaussian Process Upper Confidence Bound Achieves Nearly-Optimal Regret in Noise-Free Gaussian Process Bandits
- Improved Regret Bounds for Gaussian Process Upper Confidence Bound in Bayesian Optimization
- Improved Regret Bounds for Gaussian Process Upper Confidence Bound in Bayesian Optimization
- MALinZero: Efficient Low-Dimensional Search for Mastering Complex Multi-Agent Planning
- MAT-Agent: Adaptive Multi-Agent Training Optimization
- Tractable Multinomial Logit Contextual Bandits with Non-Linear Utilities
- Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits