NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
bandit problems
4 papers
Constrained Best Arm Identification
Greedy Algorithms for Structured Bandits: A Sharp Characterization of Asymptotic Success / Failure
Provably Efficient Multi-Task Meta Bandit Learning via Shared Representations
REINFORCE Converges to Optimal Policies with Any Learning Rate