multi-armed bandit

A problem formulation in decision theory and reinforcement learning where an agent must choose between multiple options with uncertain rewards, aiming to maximize total gain.

9 papers