On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

The stochastic gradient descent (SGD) method and its variants are algorithms of choice for many Deep Learning tasks. These methods operate in a small-batch regime wherein a fraction of the training data, say $32$-$512$ data points, is sampled to compute an approximation to the gradient. It has been observed in practice that when using a larger batch there is a degradation in the quality of the model, as measured by its ability to generalize. We investigate the cause for this generalization drop in the large-batch regime and present numerical evidence that supports the view that large-batch methods tend to converge to sharp minimizers of the training and testing functions - and as is well known, sharp minima lead to poorer generalization. In contrast, small-batch methods consistently converge to flat minimizers, and our experiments support a commonly held view that this is due to the inherent noise in the gradient estimation. We discuss several strategies to attempt to help large-batch methods eliminate this generalization gap.

Gradient-based learningapplied to document…Gradient-based learning applied to document recognitionLarge Scale DistributedDeep NetworksLarge Scale Distributed Deep NetworksOn the importance ofinitialization and…On the importance of initialization and momentum in deep learningDropout: a simple way toprevent neural networks…Dropout: a simple way to prevent neural networks from overfittingBatch Normalization:Accelerating Deep…Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate ShiftQualitativelycharacterizing neural…Qualitatively characterizing neural network optimization problemsExplaining andHarnessing Adversarial…Explaining and Harnessing Adversarial ExamplesThe Loss Surfaces ofMultilayer NetworksThe Loss Surfaces of Multilayer NetworksVery Deep ConvolutionalNetworks for Large-Scal…Very Deep Convolutional Networks for Large-Scale Image RecognitionTrain faster, generalizebetter: Stability of…Train faster, generalize better: Stability of stochastic gradient descentNo bad local minima:Data independent…No bad local minima: Data independent training error guarantees for multilayer neural networksEntropy-SGD: BiasingGradient Descent Into…Entropy-SGD: Biasing Gradient Descent Into Wide ValleysA Progressive BatchingL-BFGS Method for…A Progressive Batching L-BFGS Method for Machine LearningUnderstanding BatchNormalizationUnderstanding Batch NormalizationOptimization of neuralnetworks via…Optimization of neural networks via finite-value quantum fluctuationsSuper-Convergence: VeryFast Training of…Super-Convergence: Very Fast Training of Residual Networks Using Large Learning RatesOverparameterizedNonlinear Learning…Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?An Empirical Study ofExample Forgetting…An Empirical Study of Example Forgetting during Deep Neural Network LearningLocal SGD Converges Fastand Communicates LittleLocal SGD Converges Fast and Communicates LittleInput and Weight SpaceSmoothing for…Input and Weight Space Smoothing for Semi-Supervised LearningA Gram-Gauss-NewtonMethod Learning…A Gram-Gauss-Newton Method Learning Overparameterized Deep Neural Networks for Regression ProblemsFantastic GeneralizationMeasures and Where to…Fantastic Generalization Measures and Where to Find ThemRethinking ParameterCounting in Deep Models…Rethinking Parameter Counting in Deep Models: Effective Dimensionality RevisitedGeneralization boundsfor deep learningGeneralization bounds for deep learningOn Large-Batch Trainingfor Deep Learning…On Large-Batch Training for Deep Learning: Generalization Gap and Sharp MinimaEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.