On Exact Computation with an Infinitely Wide Neural Net

How well does a classic deep net architecture like AlexNet or VGG19 classify on a standard dataset such as CIFAR-10 when its width --- namely, number of channels in convolutional layers, and number of nodes in fully-connected internal layers --- is allowed to increase to infinity? Such questions have come to the forefront in the quest to theoretically understand deep learning and its mysteries about optimization and generalization. They also connect deep learning to notions such as Gaussian processes and kernels. A recent paper [Jacot et al., 2018] introduced the Neural Tangent Kernel (NTK) which captures the behavior of fully-connected deep nets in the infinite width limit trained by gradient descent; this object was implicit in some other recent papers. An attraction of such ideas is that a pure kernel-based method is used to capture the power of a fully-trained deep net of infinite width. The current paper gives the first efficient exact algorithm for computing the extension of NTK to convolutional neural nets, which we call Convolutional NTK (CNTK), as well as an efficient GPU implementation of this algorithm. This results in a significant new benchmark for the performance of a pure kernel-based method on CIFAR-10, being $10\%$ higher than the methods reported in [Novak et al., 2019], and only $6\%$ lower than the performance of the corresponding finite deep net architecture (once batch normalization, etc. are turned off). Theoretically, we also give the first non-asymptotic proof showing that a fully-trained sufficiently wide net is indeed equivalent to the kernel regression predictor using NTK.

SGD Learns the ConjugateKernel Class of the…SGD Learns the Conjugate Kernel Class of the NetworkNeural Tangent Kernel:Convergence and…Neural Tangent Kernel: Convergence and Generalization in Neural NetworksStochastic GradientDescent Optimizes…Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU NetworksA Convergence Theory forDeep Learning via…A Convergence Theory for Deep Learning via Over-ParameterizationGaussian ProcessBehaviour in Wide Deep…Gaussian Process Behaviour in Wide Deep Neural NetworksGradient Descent FindsGlobal Minima of Deep…Gradient Descent Finds Global Minima of Deep Neural NetworksScaling Limits of WideNeural Networks with…Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel DerivationFine-Grained Analysis ofOptimization and…Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural NetworksWide Neural Networks ofAny Depth Evolve as…Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient DescentA Note on Lazy Trainingin Supervised…A Note on Lazy Training in Supervised Differentiable ProgrammingLearning andGeneralization in…Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two LayersGeneralization ErrorBounds of Gradient…Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep ReLU NetworksWhat Can ResNet LearnEfficiently, Going…What Can ResNet Learn Efficiently, Going Beyond Kernels?Regularization Matters:Generalization and…Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced KernelEnhanced ConvolutionalNeural Tangent KernelsEnhanced Convolutional Neural Tangent KernelsLearning andGeneralization in…Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two LayersA Fine-Grained SpectralPerspective on Neural…A Fine-Grained Spectral Perspective on Neural NetworksDynamics of Deep NeuralNetworks and Neural…Dynamics of Deep Neural Networks and Neural Tangent HierarchyBeyond Linearization: OnQuadratic and…Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksLearningOver-Parametrized…Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTKTensor Programs II:Neural Tangent Kernel…Tensor Programs II: Neural Tangent Kernel for Any ArchitectureDisentanglingTrainability and…Disentangling Trainability and Generalization in Deep Neural NetworksBackward FeatureCorrection: How Deep…Backward Feature Correction: How Deep Learning Performs Deep (Hierarchical) LearningTowards UnderstandingEnsemble, Knowledge…Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningOn Exact Computationwith an Infinitely Wide…On Exact Computation with an Infinitely Wide Neural Net過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。