Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers

Christoph Lampert (Institute of Science and Technology Austria (ISTA)) · Peter Súkeník (Institute of Science and Technology Austria) · Marco Mondelli (IST Austria)
computer visioncross entropy lossdata-agnostic modelsdeep neural networksdeep regularized transformersfeature representationsglobal optimalanguage datasetslarge-depth resnetlayernormmean squared error lossmulti-layer perceptronsneural collapseresidual networkstheoretical resultsunconstrained features model

The empirical emergence of neural collapse