NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
adamw
3 papers
How to Scale Second-Order Optimization
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks