NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
gradient steps
3 papers
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
Conditioning Matters: Training Diffusion Policies is Faster Than You Think
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training