convergence rate
The speed at which an iterative algorithm approaches its solution or optimum. In AI, understanding the convergence rate can help in optimizing training time and ensuring efficient learning.
- A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
- Accelerated Vertical Federated Adversarial Learning through Decoupling Layer-Wise Dependencies
- Adaptive and Multi-scale Affinity Alignment for Hierarchical Contrastive Learning
- Advancing Wasserstein Convergence Analysis of Score-Based Models: Insights from Discretization and Second-Order Acceleration
- Any-stepsize Gradient Descent for Separable Data under Fenchel–Young Losses
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
- Convergence of Clipped SGD on Convex $(L_0,L_1)$-Smooth Functions
- Efficient Federated Learning against Byzantine Attacks and Data Heterogeneity via Aggregating Normalized Gradients
- Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
- Exploring the Noise Robustness of Online Conformal Prediction
- Feature Unlearning: Theoretical Foundations and Practical Applications with Shuffling
- Finding Low-Rank Matrix Weights in DNNs via Riemannian Optimization: RAdaGrad and RAdamW
- Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
- LLM at Network Edge: A Layer-wise Efficient Federated Fine-tuning Approach
- Learning to Reason under Off-Policy Guidance
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Mixed-Sample SGD: an End-to-end Analysis of Supervised Transfer Learning
- Nearly Dimension-Independent Convergence of Mean-Field Black-Box Variational Inference
- On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
- On the Convergence of Stochastic Smoothed Multi-Level Compositional Gradient Descent Ascent
- PaZO: Preconditioned Accelerated Zeroth-Order Optimization for Fine-Tuning LLMs
- PoLAR: Polar-Decomposed Low-Rank Adapter Representation
- Problem-Parameter-Free Decentralized Bilevel Optimization
- Semi-supervised Vertex Hunting, with Applications in Network and Text Analysis
- Tight High-Probability Bounds for Nonconvex Heavy-Tailed Scenario under Weaker Assumptions
- Tight analyses of first-order methods with error feedback
- Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach
- Uncoupled and Convergent Learning in Monotone Games under Bandit Feedback