Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks

Recent works have cast some light on the mystery of why deep nets fit any data and generalize despite being very overparametrized. This paper analyzes training and generalization for a simple 2-layer ReLU net with random initialization, and provides the following improvements over recent works: (i) Using a tighter characterization of training speed than recent papers, an explanation for why training a neural net with random labels leads to slower training, as originally observed in [Zhang et al. ICLR'17]. (ii) Generalization bound independent of network size, using a data-dependent complexity measure. Our measure distinguishes clearly between random labels and true labels on MNIST and CIFAR, as shown by experiments. Moreover, recent papers require sample complexity to increase (slowly) with the size, while our sample complexity is completely independent of the network size. (iii) Learnability of a broad class of smooth functions by 2-layer ReLU nets trained via gradient descent. The key idea is to track dynamics of training and generalization via properties of a related kernel.

Globally OptimalGradient Descent for a…Globally Optimal Gradient Descent for a ConvNet with Gaussian InputsStochastic GradientDescent Optimizes…Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU NetworksOn the Power ofOver-parametrization in…On the Power of Over-parametrization in Neural Networks with Quadratic ActivationLearningOverparameterized Neura…Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured DataNeural Networks withFinite Intrinsic…Neural Networks with Finite Intrinsic Dimension have no Spurious ValleysGradient Descent LearnsOne-hidden-layer CNN…Gradient Descent Learns One-hidden-layer CNN: Don't be Afraid of Spurious Local MinimaWhen is a ConvolutionalFilter Easy To Learn?When is a Convolutional Filter Easy To Learn?Gradient Descent FindsGlobal Minima of Deep…Gradient Descent Finds Global Minima of Deep Neural NetworksRegularization Matters:Generalization and…Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced KernelLearning andGeneralization in…Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two LayersGradient DescentProvably Optimizes…Gradient Descent Provably Optimizes Over-parameterized Neural NetworksA Note on Lazy Trainingin Supervised…A Note on Lazy Training in Supervised Differentiable ProgrammingOn Exact Computationwith an Infinitely Wide…On Exact Computation with an Infinitely Wide Neural NetAn Improved Analysis ofTraining…An Improved Analysis of Training Over-parameterized Deep Neural NetworksConvergence ofAdversarial Training in…Convergence of Adversarial Training in Overparametrized Neural NetworksA Gram-Gauss-NewtonMethod Learning…A Gram-Gauss-Newton Method Learning Overparameterized Deep Neural Networks for Regression ProblemsOn the Inductive Bias ofNeural Tangent KernelsOn the Inductive Bias of Neural Tangent KernelsBeyond Linearization: OnQuadratic and…Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksDynamics of Deep NeuralNetworks and Neural…Dynamics of Deep Neural Networks and Neural Tangent HierarchyQuantifying thegeneralization error in…Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothnessGeneralized LeverageScore Sampling for…Generalized Leverage Score Sampling for Neural NetworksTrainingOver-parameterized Deep…Training Over-parameterized Deep ResNet Is almost as Easy as Training a Two-layer NetworkLinearized two-layersneural networks in high…Linearized two-layers neural networks in high dimensionDemystifying the GlobalConvergence Puzzle of…Demystifying the Global Convergence Puzzle of Learning Over-parameterized ReLU Nets in Very High DimensionsFine-Grained Analysis ofOptimization and…Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural NetworksEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.