NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Andrew Wilson
3 papers
New York University
How to Scale Second-Order Optimization
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is Wasteful
Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion