sgd
- Flat Channels to Infinity in Neural Loss Landscapes
- Momentum-SAM: Sharpness Aware Minimization without Computational Overhead
- On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
- On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization