Denny Wu
- Emergence and scaling laws in SGD learning of shallow neural networks
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
- Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
- When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective