Non-Singularity of the Gradient Descent Map for Neural Networks with Piecewise Analytic Activations
convergence analysisconvolutional layersfully connected layersglobal minimagradient descentlearning dynamicslipschitz smoothnessloss landscapeneural network architecturesnon-singular mappiecewise analytic activationssaddle pointssoftmax attentionstep-sizesstochastic gd
The theory of training deep networks has become a central question of modern machine learning and has inspired many practical advancements. In particular, the gradient descent (GD) optimization algorithm has been extensively studied in recent years. A key assumption about GD has appeared in several recent works: the \emph{GD map is non-singular}