training dynamics
The evolution of the learning process over time, including how model parameters change and how performance metrics improve during the training phase.
- A Minimalist Example of Edge-of-Stability and Progressive Sharpening
- ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
- Any-stepsize Gradient Descent for Separable Data under Fenchel–Young Losses
- Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
- Boosting Resilience of Large Language Models through Causality-Driven Robust Optimization
- Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
- Composing Global Solutions to Reasoning Tasks via Algebraic Objects in Neural Nets
- Contrastive Learning with Data Misalignment: Feature Purity, Training Dynamics and Theoretical Generalization Guarantees
- Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
- Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
- DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
- Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
- Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
- EvoLM: In Search of Lost Language Model Training Dynamics
- EvoLM: In Search of Lost Language Model Training Dynamics
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
- Generative Data Augmentation via Diffusion Distillation, Adversarial Alignment, and Importance Reweighting
- Global Minimizers of Sigmoid Contrastive Loss
- Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks
- Impact of Layer Norm on Memorization and Generalization in Transformers
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Learning to Add, Multiply, and Execute Algorithmic Instructions Exactly with Neural Networks
- Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
- Less is More: Local Intrinsic Dimensions of Contextual Language Models
- Memorization in Graph Neural Networks
- NTKMTL: Mitigating Task Imbalance in Multi-Task Learning from Neural Tangent Kernel Perspective
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- PALQO: Physics-informed model for Accelerating Large-scale Quantum Optimization
- Quantitative convergence of trained neural networks to Gaussian processes
- Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
- Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
- Scaling Law with Learning Rate Annealing
- Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
- Self-Assembling Graph Perceptrons
- Structure Matters: Dynamic Policy Gradient
- TITAN: A Trajectory-Informed Technique for Adaptive Parameter Freezing in Large-Scale VQE
- The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis
- Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
- Understanding the Evolution of the Neural Tangent Kernel at the Edge of Stability
- When majority rules, minority loses: bias amplification of gradient descent
- Why Diffusion Models Don’t Memorize: The Role of Implicit Dynamical Regularization in Training
- Why Diffusion Models Don’t Memorize: The Role of Implicit Dynamical Regularization in Training