Exact solutions to the nonlinear dynamics of learning in deep linear neural networks

Abstract: Despite the widespread practical success of deep learning methods, our theoretical understanding of the dynamics of learning in deep neural networks remains quite sparse. We attempt to bridge the gap between the theory and practice of deep learning by systematically analyzing learning dynamics for the restricted case of deep linear neural networks. Despite the linearity of their input-output map, such networks have nonlinear gradient descent dynamics on weights that change with the addition of each new hidden layer. We show that deep linear networks exhibit nonlinear learning phenomena similar to those seen in simulations of nonlinear networks, including long plateaus followed by rapid transitions to lower error solutions, and faster convergence from greedy unsupervised pretraining initial conditions than from random initial conditions. We provide an analytical description of these phenomena by finding new exact solutions to the nonlinear dynamics of deep learning. Our theoretical analysis also reveals the surprising finding that as the depth of a network approaches infinity, learning speed can nevertheless remain finite: for a special class of initial conditions on the weights, very deep networks incur only a finite, depth independent, delay in learning speed relative to shallow networks. We show that, under certain conditions on the training data, unsupervised pretraining can find this special class of initial conditions, while scaled random Gaussian initializations cannot. We further exhibit a new class of random orthogonal initial conditions on weights that, like unsupervised pre-training, enjoys depth independent learning times. We further show that these initial conditions also lead to faithful propagation of gradients even in deep nonlinear networks, as long as they operate in a special regime known as the edge of chaos.

Learning long-termdependencies with…Learning long-term dependencies with gradient descent is difficultGreedy Layer-WiseTraining of Deep…Greedy Layer-Wise Training of Deep NetworksReducing theDimensionality of Data…Reducing the Dimensionality of Data with Neural NetworksThe Difficulty ofTraining Deep…The Difficulty of Training Deep Architectures and the Effect of Unsupervised Pre-TrainingLearning DeepArchitectures for AILearning Deep Architectures for AIUnderstanding thedifficulty of training…Understanding the difficulty of training deep feedforward neural networksWhy Does UnsupervisedPre-training Help Deep…Why Does Unsupervised Pre-training Help Deep Learning?Deep learning viaHessian-free…Deep learning via Hessian-free optimizationImageNet Classificationwith Deep Convolutional…ImageNet Classification with Deep Convolutional Neural NetworksMulti-column Deep NeuralNetworks for Image…Multi-column Deep Neural Networks for Image ClassificationOn the importance ofinitialization and…On the importance of initialization and momentum in deep learningOn the difficulty oftraining recurrent…On the difficulty of training recurrent neural networksThe Loss Surfaces ofMultilayer NetworksThe Loss Surfaces of Multilayer NetworksHighway NetworksHighway NetworksData-Efficient Learningof Feedback Policies…Data-Efficient Learning of Feedback Policies from Image Pixels using Deep Dynamical ModelsDeep ConvolutionalNeural Networks for…Deep Convolutional Neural Networks for Image Classification: A Comprehensive ReviewBIER - BoostingIndependent Embeddings…BIER - Boosting Independent Embeddings RobustlyVehicle classificationfrom low-frequency GPS…Vehicle classification from low-frequency GPS data with recurrent neural networksExploiting artificialintelligence for…Exploiting artificial intelligence for digitally enriched museum visitsFixup Initialization:Residual Learning…Fixup Initialization: Residual Learning Without NormalizationSpectrum Concentrationin Deep Residual…Spectrum Concentration in Deep Residual Learning: A Free Probability ApproachApplication ofConvolutional Recurrent…Application of Convolutional Recurrent Neural Network for Individual Recognition Based on Resting State fMRI DataUnderstanding theDifficulty of Training…Understanding the Difficulty of Training TransformersDeep Subspace ClusteringDeep Subspace ClusteringExact solutions to thenonlinear dynamics of…Exact solutions to the nonlinear dynamics of learning in deep linear neural networks過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。