Theoretical Analysis of Auto Rate-Tuning by Batch Normalization

Batch Normalization (BN) has become a cornerstone of deep learning across diverse architectures, appearing to help optimization as well as generalization. While the idea makes intuitive sense, theoretical analysis of its effectiveness has been lacking. Here theoretical support is provided for one of its conjectured properties, namely, the ability to allow gradient descent to succeed with less tuning of learning rates. It is shown that even if we fix the learning rate of scale-invariant parameters (e.g., weights of each layer with BN) to a constant (say, $0.3$), gradient descent still approaches a stationary point (i.e., a solution where gradient is zero) in the rate of $T^{-1/2}$ in $T$ iterations, asymptotically matching the best bound for gradient descent with well-tuned learning rates. A similar result with convergence rate $T^{-1/4}$ is also shown for stochastic gradient descent.

Adam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationVery Deep ConvolutionalNetworks for Large-Scal…Very Deep Convolutional Networks for Large-Scale Image RecognitionDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionLayer NormalizationLayer NormalizationWeight Normalization: ASimple…Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural NetworksInception-v4,Inception-ResNet and th…Inception-v4, Inception-ResNet and the Impact of Residual Connections on LearningWNGrad: Learn theLearning Rate in…WNGrad: Learn the Learning Rate in Gradient DescentGroup NormalizationGroup NormalizationExponential convergencerates for Batch…Exponential convergence rates for Batch Normalization: The power of length-direction decoupling in non-convex optimizationAdaGrad stepsizes: Sharpconvergence over…AdaGrad stepsizes: Sharp convergence over nonconvex landscapes, from any initializationA Unified Analysis ofAdaGrad With Weighted…A Unified Analysis of AdaGrad With Weighted Aggregation and Momentum AccelerationOn the Convergence ofAdaptive Gradient…On the Convergence of Adaptive Gradient Methods for Nonconvex OptimizationSwitchable Normalizationfor…Switchable Normalization for Learning-to-Normalize Deep RepresentationLuck Matters:Understanding Training…Luck Matters: Understanding Training Dynamics of Deep ReLU NetworksNew Interpretations ofNormalization Methods i…New Interpretations of Normalization Methods in Deep LearningImplicit Regularizationand Convergence for…Implicit Regularization and Convergence for Weight NormalizationOptimization for DeepLearning: An OverviewOptimization for Deep Learning: An OverviewBatch NormalizationBiases Residual Blocks…Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksRevisiting InternalCovariate Shift for…Revisiting Internal Covariate Shift for Batch NormalizationAn Investigation Intothe Stochasticity of…An Investigation Into the Stochasticity of Batch WhiteningGraphNorm: A PrincipledApproach to Acceleratin…GraphNorm: A Principled Approach to Accelerating Graph Neural Network TrainingNoether's LearningDynamics: The Role of…Noether's Learning Dynamics: The Role of Kinetic Symmetry Breaking in Deep LearningNormalization Techniquesin Training DNNs…Normalization Techniques in Training DNNs: Methodology, Analysis and ApplicationTransformers withoutNormalizationTransformers without NormalizationTheoretical Analysis ofAuto Rate-Tuning by…Theoretical Analysis of Auto Rate-Tuning by Batch Normalization過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。