kl divergence
Kullback-Leibler divergence is a measure of how one probability distribution diverges from a second, expected probability distribution. In AI, it is often used in optimization problems, particularly within variational inference and generative models to quantify the difference between learned representations and true distributions.
- $Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
- Absorb and Converge: Provable Convergence Guarantee for Absorbing Discrete Diffusion Models
- AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
- Beyond Scores: Proximal Diffusion Models
- Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
- Continuous-time Riemannian SGD and SVRG Flows on Wasserstein Probabilistic Space
- Discrete Diffusion Models: Novel Analysis and New Sampler Guarantees
- Distributional LLM-as-a-Judge
- Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms
- Foundations of Top-$k$ Decoding for Language Models
- Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
- Instance-Optimality for Private KL Distribution Estimation
- LLM Safety Alignment is Divergence Estimation in Disguise
- MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework
- Missing Data Imputation by Reducing Mutual Information with Rectified Flows
- Non-convex entropic mean-field optimization via Best Response flow
- On Extending Direct Preference Optimization to Accommodate Ties
- PLD: A Choice-Theoretic List-Wise Knowledge Distillation
- PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation
- Preference Distillation via Value based Reinforcement Learning
- RankMatch: A Novel Approach to Semi-Supervised Label Distribution Learning Leveraging Rank Correlation between Labels
- Risk-aware Direct Preference Optimization under Nested Risk Measure
- Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
- Simple Distillation for One-Step Diffusion Models
- SimpleStrat: Diversifying Language Model Generation with Stratification
- Solving Discrete (Semi) Unbalanced Optimal Transport with Equivalent Transformation Mechanism and KKT-Multiplier Regularization