Sharp Minima Can Generalize For Deep Nets

Despite their overwhelming capacity to overfit, deep learning architectures tend to generalize relatively well to unseen data, allowing them to be deployed in practice. However, explaining why this is the case is still an open area of research. One standing hypothesis that is gaining popularity, e.g. Hochreiter & Schmidhuber (1997); Keskar et al. (2017), is that the flatness of minima of the loss function found by stochastic gradient based methods results in good generalization. This paper argues that most notions of flatness are problematic for deep models and can not be directly applied to explain generalization. Specifically, when focusing on deep networks with rectifier units, we can exploit the particular geometry of parameter space induced by the inherent symmetries that these architectures exhibit to build equivalent models corresponding to arbitrarily sharper minima. Furthermore, if we allow to reparametrize a function, the geometry of its parameters can change drastically without affecting its generalization properties.

Path-SGD:Path-Normalized…Path-SGD: Path-Normalized Optimization in Deep Neural NetworksThe Loss Surfaces ofMultilayer NetworksThe Loss Surfaces of Multilayer NetworksBatch Normalization:Accelerating Deep…Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate ShiftExplaining andHarnessing Adversarial…Explaining and Harnessing Adversarial ExamplesDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionLocal minima in trainingof deep networksLocal minima in training of deep networksWeight Normalization: ASimple…Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural NetworksImproved Techniques forTraining GANsImproved Techniques for Training GANsOn Large-Batch Trainingfor Deep Learning…On Large-Batch Training for Deep Learning: Generalization Gap and Sharp MinimaEntropy-SGD: BiasingGradient Descent Into…Entropy-SGD: Biasing Gradient Descent Into Wide ValleysUnderstanding deeplearning requires…Understanding deep learning requires rethinking generalizationDensity estimation usingReal NVPDensity estimation using Real NVPFisher-Rao Metric,Geometry, and Complexit…Fisher-Rao Metric, Geometry, and Complexity of Neural NetworksEmpirical Analysis ofthe Hessian of…Empirical Analysis of the Hessian of Over-Parametrized Neural NetworksLearningOverparameterized Neura…Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured DataAn Investigation intoNeural Net Optimization…An Investigation into Neural Net Optimization via Hessian Eigenvalue DensityOn the Relation Betweenthe Sharpest Directions…On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step LengthOrthogonal Deep NeuralNetworksOrthogonal Deep Neural NetworksEmergent properties ofthe local geometry of…Emergent properties of the local geometry of neural loss landscapesFantastic GeneralizationMeasures and Where to…Fantastic Generalization Measures and Where to Find ThemBayesian Deep Learningand a Probabilistic…Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationTraces ofClass/Cross-Class…Traces of Class/Cross-Class Structure Pervade Deep Learning SpectraQuantifying thegeneralization error in…Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothnessSharpness-AwareMinimization for…Sharpness-Aware Minimization for Efficiently Improving GeneralizationSharp Minima CanGeneralize For Deep NetsSharp Minima Can Generalize For Deep Nets過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。