Shinji Ito
- Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
- Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
- Optimal Regret of Bandits under Differential Privacy
- Revisiting Follow-the-Perturbed-Leader with Unbounded Perturbations in Bandit Problems