Just One Layer Norm Guarantees Stable Extrapolation
bounded-variance kernelconvergenceempirical experimentsextrapolationextrapolatory stabilityfacial age estimationfinite-width networksinfinitely-wide neural networkslayer normneural tangent kernelpathologically large outputsresidue size predictiontheoretical findingstraining distributionuncontrolled growthunderrepresented ethnicities
In spite of their prevalence, the behaviour of Neural Networks when extrapolating far from the training distribution remains poorly understood, with existing results limited to specific cases. In this work, we prove general results