The Low-Rank Simplicity Bias in Deep Networks

Modern deep neural networks are highly over-parameterized compared to the data on which they are trained, yet they often generalize remarkably well. A flurry of recent work has asked: why do deep networks not overfit to their training data? In this work, we make a series of empirical observations that investigate and extend the hypothesis that deeper networks are inductively biased to find solutions with lower effective rank embeddings. We conjecture that this bias exists because the volume of functions that maps to low effective rank embedding increases with depth. We show empirically that our claim holds true on finite width linear and non-linear models on practical learning paradigms and show that on natural data, these are often the solutions that generalize well. We then show that the simplicity bias exists at both initialization and after training and is resilient to hyper-parameters and learning methods. We further demonstrate how linear over-parameterization of deep non-linear models can be used to induce low-rank bias, improving generalization performance on CIFAR and ImageNet without changing the modeling capacity.

ImageNet Classificationwith Deep Convolutional…ImageNet Classification with Deep Convolutional Neural NetworksAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationBatch Normalization:Accelerating Deep…Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate ShiftGoing Deeper withConvolutionsGoing Deeper with ConvolutionsImageNet Large ScaleVisual Recognition…ImageNet Large Scale Visual Recognition ChallengeDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionAggregated ResidualTransformations for Dee…Aggregated Residual Transformations for Deep Neural NetworksSGD on Neural NetworksLearns Functions of…SGD on Neural Networks Learns Functions of Increasing ComplexityScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsDeep Double Descent:Where Bigger Models and…Deep Double Descent: Where Bigger Models and More Data HurtGradient Starvation: ALearning Proclivity in…Gradient Starvation: A Learning Proclivity in Neural NetworksAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleCan contrastive learningavoid shortcut…Can contrastive learning avoid shortcut solutions?Indiscriminate PoisoningAttacks Are ShortcutsIndiscriminate Poisoning Attacks Are ShortcutsOvercoming The SpectralBias of Neural Value…Overcoming The Spectral Bias of Neural Value ApproximationNeural Fields in VisualComputing and BeyondNeural Fields in Visual Computing and BeyondThe Law of Parsimony inGradient Descent for…The Law of Parsimony in Gradient Descent for Learning Deep Linear NetworksStochastic Collapse: HowGradient Noise Attracts…Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksMachine learning anddeep learning—A review…Machine learning and deep learning—A review for ecologistsAttribute-Aware DeepHashing With…Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image RetrievalSimplicity Bias in1-Hidden Layer Neural…Simplicity Bias in 1-Hidden Layer Neural NetworksOutliers with OpposingSignals Have an Outsize…Outliers with Opposing Signals Have an Outsized Effect on Neural Network OptimizationNeural Redshift: RandomNetworks are not Random…Neural Redshift: Random Networks are not Random FunctionsDo Neural Networks NeedGradient Descent to…Do Neural Networks Need Gradient Descent to Generalize? A Theoretical StudyThe Low-Rank SimplicityBias in Deep NetworksThe Low-Rank Simplicity Bias in Deep Networks過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。