upper confidence bound

A strategy in bandit problems where the algorithm calculates an upper confidence limit for the expected rewards of each action, balancing exploration and exploitation.

7 papers