Implicit Regularization in Deep Matrix Factorization

Efforts to understand the generalization mystery in deep learning have led to the belief that gradient-based optimization induces a form of implicit regularization, a bias towards models of low "complexity." We study the implicit regularization of gradient descent over deep linear neural networks for matrix completion and sensing, a model referred to as deep matrix factorization. Our first finding, supported by theory and experiments, is that adding depth to a matrix factorization enhances an implicit tendency towards low-rank solutions, oftentimes leading to more accurate recovery. Secondly, we present theoretical and empirical arguments questioning a nascent view by which implicit regularization in matrix factorization can be captured using simple mathematical norms. Our results point to the possibility that the language of standard regularizers may not be rich enough to fully encompass the implicit regularization brought forth by gradient-based optimization.

Guaranteed Minimum-RankSolutions of Linear…Guaranteed Minimum-Rank Solutions of Linear Matrix Equations via Nuclear Norm MinimizationThe MovieLens Datasets:History and ContextThe MovieLens Datasets: History and ContextIn Search of the RealInductive Bias: On the…In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep LearningAn Overview of Low-RankMatrix Recovery From…An Overview of Low-Rank Matrix Recovery From Incomplete ObservationsNo Spurious Local Minimain Nonconvex Low Rank…No Spurious Local Minima in Nonconvex Low Rank Problems: A Unified Geometric AnalysisAutomaticdifferentiation in…Automatic differentiation in PyTorchIdentity Matters in DeepLearningIdentity Matters in Deep LearningAlgorithmicRegularization in…Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically BalancedWidth Provably Mattersin Optimization for Dee…Width Provably Matters in Optimization for Deep Linear Neural NetworksImplicit Regularizationof Discrete Gradient…Implicit Regularization of Discrete Gradient Dynamics in Deep Linear Neural NetworksOn the Spectral Bias ofDeep Neural NetworksOn the Spectral Bias of Deep Neural NetworksNonconvex OptimizationMeets Low-Rank Matrix…Nonconvex Optimization Meets Low-Rank Matrix Factorization: An OverviewTraining Linear NeuralNetworks: Non-Local…Training Linear Neural Networks: Non-Local Convergence and Complexity ResultsKernel and Rich Regimesin Overparametrized…Kernel and Rich Regimes in Overparametrized ModelsDeep matrixfactorizationsDeep matrix factorizationsUnderstanding ImplicitRegularization in…Understanding Implicit Regularization in Over-Parameterized Nonlinear Statistical ModelUnderstanding deeplearning (still)…Understanding deep learning (still) requires rethinking generalizationA survey on deep matrixfactorizationsA survey on deep matrix factorizationsRevealing the Structureof Deep Neural Networks…Revealing the Structure of Deep Neural Networks via Convex DualityAn UnconstrainedLayer-Peeled Perspectiv…An Unconstrained Layer-Peeled Perspective on Neural CollapseUnderstandingDimensional Collapse in…Understanding Dimensional Collapse in Contrastive Self-supervised LearningWhat Happens after SGDReaches Zero Loss? -A…What Happens after SGD Reaches Zero Loss? -A Mathematical FrameworkStochastic Training isNot Necessary for…Stochastic Training is Not Necessary for GeneralizationThe ImplicitRegularization of…The Implicit Regularization of Momentum Gradient Descent in Overparametrized ModelsImplicit Regularizationin Deep Matrix…Implicit Regularization in Deep Matrix FactorizationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.