Prevalence of Neural Collapse during the terminal phase of deep learning training

Modern practice for training classification deepnets involves a Terminal Phase of Training (TPT), which begins at the epoch where training error first vanishes; During TPT, the training error stays effectively zero while training loss is pushed towards zero. Direct measurements of TPT, for three prototypical deepnet architectures and across seven canonical classification datasets, expose a pervasive inductive bias we call Neural Collapse, involving four deeply interconnected phenomena: (NC1) Cross-example within-class variability of last-layer training activations collapses to zero, as the individual activations themselves collapse to their class-means; (NC2) The class-means collapse to the vertices of a Simplex Equiangular Tight Frame (ETF); (NC3) Up to rescaling, the last-layer classifiers collapse to the class-means, or in other words to the Simplex ETF, i.e. to a self-dual configuration; (NC4) For a given activation, the classifier's decision collapses to simply choosing whichever class has the closest train class-mean, i.e. the Nearest Class Center (NCC) decision rule. The symmetric and very simple geometry induced by the TPT confers important benefits, including better generalization performance, better robustness, and better interpretability.

THE USE OF MULTIPLEMEASUREMENTS IN…THE USE OF MULTIPLE MEASUREMENTS IN TAXONOMIC PROBLEMSProceedings of IEEEConference on Computer…Proceedings of IEEE Conference on Computer Vision and Pattern RecognitionAdvances in neuralinformation processing…Advances in neural information processing systems 7Proceedings of the 24thinternational conferenc…Proceedings of the 24th international conference on Machine learningImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseVery Deep ConvolutionalNetworks for Large-Scal…Very Deep Convolutional Networks for Large-Scale Image RecognitionDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionThe Full Spectrum ofDeep Net Hessians At…The Full Spectrum of Deep Net Hessians At Scale: Dynamics with Sample SizeReconciling modernmachine-learning…Reconciling modern machine-learning practice and the classical bias–variance trade-offMeasurements ofThree-Level Hierarchica…Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet HessiansPyTorch: An ImperativeStyle, High-Performance…PyTorch: An Imperative Style, High-Performance Deep Learning LibraryLearning Multiple Layersof Features from Tiny…Learning Multiple Layers of Features from Tiny ImagesWhy Do Better LossFunctions Lead to Less…Why Do Better Loss Functions Lead to Less Transferable Features?Phase Collapse in NeuralNetworksPhase Collapse in Neural NetworksNeural collapse withunconstrained featuresNeural collapse with unconstrained featuresReduNet: A White-boxDeep Network from the…ReduNet: A White-box Deep Network from the Principle of Maximizing Rate ReductionBalanced ContrastiveLearning for Long-Taile…Balanced Contrastive Learning for Long-Tailed Visual RecognitionImage2Point: 3DPoint-Cloud…Image2Point: 3D Point-Cloud Understanding with 2D Image Pretrained ModelsAsymmetric LossFunctions for…Asymmetric Loss Functions for Noise-Tolerant Learning: Theory and ApplicationsGEN: Pushing the Limitsof Softmax-Based…GEN: Pushing the Limits of Softmax-Based Out-of-Distribution DetectionDeep neural networksarchitectures from the…Deep neural networks architectures from the perspective of manifold learningGaussian universality ofperceptrons with random…Gaussian universality of perceptrons with random labelsFederated deeplong-tailed learning: A…Federated deep long-tailed learning: A surveyConditional MutualInformation Constrained…Conditional Mutual Information Constrained Deep Learning: Framework and Preliminary ResultsPrevalence of NeuralCollapse during the…Prevalence of Neural Collapse during the terminal phase of deep learning trainingEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.