Peng Zhao
- Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update
- Gradient-Variation Online Adaptivity for Accelerated Optimization with Hölder Smoothness
- Learning Memory-Enhanced Improvement Heuristics for Flexible Job Shop Scheduling
- Optimistic Online-to-Batch Conversions for Accelerated Convergence and Universality
- Parameter-free Algorithms for the Stochastically Extended Adversarial Model
- Provably Efficient Online RLHF with One-Pass Reward Modeling