stochastic environments
- Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
- Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
- Offline Goal-conditioned Reinforcement Learning with Quasimetric Representations
- PlanU: Large Language Model Reasoning through Planning under Uncertainty