A Geometric Analysis of Neural Collapse with Unconstrained Features

We provide the first global optimization landscape analysis of $Neural\;Collapse$ -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training. As recently reported by Papyan et al., this phenomenon implies that ($i$) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and ($ii$) cross-example within-class variability of last-layer activations collapses to zero. We study the problem based on a simplified $unconstrained\;feature\;model$, which isolates the topmost layers from the classifier of the neural network. In this context, we show that the classical cross-entropy loss with weight decay has a benign global landscape, in the sense that the only global minimizers are the Simplex ETFs while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. In contrast to existing landscape analysis for deep neural networks which is often disconnected from practice, our analysis of the simplified model not only does it explain what kind of features are learned in the last layer, but it also shows why they can be efficiently optimized in the simplified settings, matching the empirical observations in practical deep network architectures. These findings could have profound implications for optimization, generalization, and robustness of broad interests. For example, our experiments demonstrate that one may set the feature dimension equal to the number of classes and fix the last-layer classifier to be a Simplex ETF for network training, which reduces memory cost by over $20\%$ on ResNet18 without sacrificing the generalization performance.

Adam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationBatch Normalization:Accelerating Deep…Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate ShiftDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionMatrix Completion has NoSpurious Local MinimumMatrix Completion has No Spurious Local MinimumGlobal Optimality inLow-Rank Matrix…Global Optimality in Low-Rank Matrix OptimizationUnderstanding deeplearning requires…Understanding deep learning requires rethinking generalizationThe non-convex geometryof low-rank matrix…The non-convex geometry of low-rank matrix optimizationNeural Collapse withCross-Entropy LossNeural Collapse with Cross-Entropy LossPrevalence of NeuralCollapse during the…Prevalence of Neural Collapse during the terminal phase of deep learning trainingA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsFrom Symmetry toGeometry: Tractable…From Symmetry to Geometry: Tractable Nonconvex ProblemsNeural collapse withunconstrained featuresNeural collapse with unconstrained featuresFrom Symmetry toGeometry: Tractable…From Symmetry to Geometry: Tractable Nonconvex ProblemsAn UnconstrainedLayer-Peeled Perspectiv…An Unconstrained Layer-Peeled Perspective on Neural CollapseNeural Collapse UnderMSE Loss: Proximity to…Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathReduNet: A White-boxDeep Network from the…ReduNet: A White-box Deep Network from the Principle of Maximizing Rate ReductionLimitations of NeuralCollapse for…Limitations of Neural Collapse for Understanding Generalization in Deep LearningUnderstanding ImbalancedSemantic Segmentation…Understanding Imbalanced Semantic Segmentation Through Neural CollapseDecoupling MaxLogit forOut-of-Distribution…Decoupling MaxLogit for Out-of-Distribution DetectionNo Fear of ClassifierBiases: Neural Collapse…No Fear of Classifier Biases: Neural Collapse Inspired Federated Learning with Synthetic and Fixed ClassifierFG-UAP:Feature-Gathering…FG-UAP: Feature-Gathering Universal Adversarial PerturbationCoReS: CompatibleRepresentations via…CoReS: Compatible Representations via StationarityLinguistic Collapse:Neural Collapse in…Linguistic Collapse: Neural Collapse in (Large) Language ModelsLearning Equi-AngularRepresentations for…Learning Equi-Angular Representations for Online Continual LearningA Geometric Analysis ofNeural Collapse with…A Geometric Analysis of Neural Collapse with Unconstrained Features過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。