finite-horizon
Finite-horizon refers to scenarios in reinforcement learning where decision-making is constrained to a fixed number of time steps. This contrasts with infinite-horizon problems that consider long-term rewards indefinitely, and influences how agents plan and evaluate actions over time.
- Achieving $\tilde{\mathcal{O}}(1/N)$ Optimality Gap in Restless Bandits through Gaussian Approximation
- Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
- Breaking the Order Barrier: Off-Policy Evaluation for Confounded POMDPs
- Deep learning for continuous-time stochastic control with jumps
- Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
- REINFORCE Converges to Optimal Policies with Any Learning Rate
- SOMBRL: Scalable and Optimistic Model-Based RL