Exploration Exploitation
concepts · 3 notes linked
Related: Reinforcement Learning · Multi Armed Bandit · Upper Confidence Bound · Deepmind · Openai · Google Brain · Reward Function Design · Movielens
Notes
- Deep Reinforcement Learning Doesn't Work Yet — Systematic critique of deep RL limitations and failure modes
- Multi-Armed Bandits in Python: Epsilon Greedy, UCB1, Bayesian UCB, and EXP3 — Four bandit algorithm implementations evaluated on MovieLens recommendation task
- Optimism in the Face of Uncertainty: the UCB1 Algorithm — Mathematical derivation and implementation of the UCB1 bandit algorithm