On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization

Conventional wisdom in deep learning states that increasing depth improves expressiveness but complicates optimization. This paper suggests that, sometimes, increasing depth can speed up optimization. The effect of depth on optimization is decoupled from expressiveness by focusing on settings where additional layers amount to overparameterization - linear neural networks, a well-studied model. Theoretical analysis, as well as experiments, show that here depth acts as a preconditioner which may accelerate convergence. Even on simple convex problems such as linear regression with $\ell_p$ loss, $p>2$, gradient descent can benefit from transitioning to a non-convex overparameterized objective, more than it would from some common acceleration schemes. We also prove that it is mathematically impossible to obtain the acceleration effect of overparametrization via gradients of any regularizer.

Adaptive SubgradientMethods for Online…Adaptive Subgradient Methods for Online Learning and Stochastic OptimizationAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationDropout: a simple way toprevent neural networks…Dropout: a simple way to prevent neural networks from overfittingExact solutions to thenonlinear dynamics of…Exact solutions to the nonlinear dynamics of learning in deep linear neural networksA Differential Equationfor Modeling Nesterov's…A Differential Equation for Modeling Nesterov's Accelerated Gradient Method: Theory and InsightsGlobal Optimality inTensor Factorization…Global Optimality in Tensor Factorization, Deep Learning, and BeyondGeneralization Boundsfor Neural Networks…Generalization Bounds for Neural Networks through Tensor FactorizationThe Loss Surfaces ofMultilayer NetworksThe Loss Surfaces of Multilayer NetworksNo bad local minima:Data independent…No bad local minima: Data independent training error guarantees for multilayer neural networksDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionIdentity Matters in DeepLearningIdentity Matters in Deep LearningSpurious Local Minimaare Common in Two-Layer…Spurious Local Minima are Common in Two-Layer ReLU Neural NetworksAlgorithmicRegularization in…Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically BalancedStochastic GradientDescent Optimizes…Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU NetworksWidth Provably Mattersin Optimization for Dee…Width Provably Matters in Optimization for Deep Linear Neural NetworksGradient DescentProvably Optimizes…Gradient Descent Provably Optimizes Over-parameterized Neural NetworksRegularization Matters:Generalization and…Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced KernelTowards moderateoverparameterization…Towards moderate overparameterization: global convergence guarantees for training shallow neural networksOverparameterizedNonlinear Learning…Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?The Impact of NeuralNetwork…The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentTraining Linear NeuralNetworks: Non-Local…Training Linear Neural Networks: Non-Local Convergence and Complexity ResultsGeneralization ErrorBounds of Gradient…Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep ReLU NetworksNeural Mechanics:Symmetry and Broken…Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsEffects of Depth, Width,and Initialization: A…Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural NetworksOn the Optimization ofDeep Networks: Implicit…On the Optimization of Deep Networks: Implicit Acceleration by OverparameterizationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.