Deep learning generalizes because the parameter-function map is biased towards simple functions

Deep neural networks (DNNs) generalize remarkably well without explicit regularization even in the strongly over-parametrized regime where classical learning theory would instead predict that they would severely overfit. While many proposals for some kind of implicit regularization have been made to rationalise this success, there is no consensus for the fundamental reason why DNNs do not strongly overfit. In this paper, we provide a new explanation. By applying a very general probability-complexity bound recently derived from algorithmic information theory (AIT), we argue that the parameter-function map of many DNNs should be exponentially biased towards simple functions. We then provide clear evidence for this strong simplicity bias in a model DNN for Boolean functions, as well as in much larger fully connected and convolutional networks applied to CIFAR10 and MNIST. As the target functions in many real problems are expected to be highly structured, this intrinsic simplicity bias helps explain why deep networks generalize well on real world problems. This picture also facilitates a novel PAC-Bayes approach where the prior is taken over the DNN input-output function space, rather than the more conventional prior over parameter space. If we assume that the training algorithm samples parameters close to uniformly within the zero-error region then the PAC-Bayes theorem can be used to guarantee good expected generalization for target functions producing high-likelihood training sets. By exploiting recently discovered connections between DNNs and Gaussian processes to estimate the marginal likelihood, we produce relatively tight generalization PAC-Bayes error bounds which correlate well with the true error on realistic datasets such as MNIST and CIFAR10 and for architectures including convolutional and fully connected networks.

Gradient-based learningapplied to document…Gradient-based learning applied to document recognitionDropout: a simple way toprevent neural networks…Dropout: a simple way to prevent neural networks from overfittingDeep learningDeep learningExplaining andHarnessing Adversarial…Explaining and Harnessing Adversarial ExamplesTowards UnderstandingGeneralization of Deep…Towards Understanding Generalization of Deep Learning: Perspective of Loss LandscapesOn Large-Batch Trainingfor Deep Learning…On Large-Batch Training for Deep Learning: Generalization Gap and Sharp MinimaInput–output maps arestrongly biased towards…Input–output maps are strongly biased towards simple outputsA Bayesian Perspectiveon Generalization and…A Bayesian Perspective on Generalization and Stochastic Gradient DescentDeep ConvolutionalNetworks as shallow…Deep Convolutional Networks as shallow Gaussian ProcessesGaussian ProcessBehaviour in Wide Deep…Gaussian Process Behaviour in Wide Deep Neural NetworksBayesian ConvolutionalNeural Networks with…Bayesian Convolutional Neural Networks with Many Channels are Gaussian ProcessesGeneralization with DeepLearning: For…Generalization with Deep Learning: For Improvement on Sensing CapabilityNeural networks are apriori biased towards…Neural networks are a priori biased towards Boolean functions with low entropyA Fine-Grained SpectralPerspective on Neural…A Fine-Grained Spectral Perspective on Neural NetworksTensor Programs I: WideFeedforward or Recurren…Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian ProcessesExplicitizing anImplicit Bias of the…Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural NetworksGeneralization boundsfor deep learningGeneralization bounds for deep learningCompression based boundfor non-compressed…Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkWhy Flatness CorrelatesWith Generalization For…Why Flatness Correlates With Generalization For Deep Neural NetworksGradient Starvation: ALearning Proclivity in…Gradient Starvation: A Learning Proclivity in Neural NetworksDo deep neural networkshave an inbuilt Occam's…Do deep neural networks have an inbuilt Occam's razor?Simplicity Bias inTransformers and their…Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean FunctionsDouble-descent curves inneural networks: a new…Double-descent curves in neural networks: a new perspective using Gaussian processesNeural Redshift: RandomNetworks are not Random…Neural Redshift: Random Networks are not Random FunctionsDeep learninggeneralizes because the…Deep learning generalizes because the parameter-function map is biased towards simple functions過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。