What Happens after SGD Reaches Zero Loss? -A Mathematical Framework

Understanding the implicit bias of Stochastic Gradient Descent (SGD) is one of the key challenges in deep learning, especially for overparametrized models, where the local minimizers of the loss function $L$ can form a manifold. Intuitively, with a sufficiently small learning rate $η$, SGD tracks Gradient Descent (GD) until it gets close to such manifold, where the gradient noise prevents further convergence. In such a regime, Blanc et al. (2020) proved that SGD with label noise locally decreases a regularizer-like term, the sharpness of loss, $\mathrm{tr}[\nabla^2 L]$. The current paper gives a general framework for such analysis by adapting ideas from Katzenberger (1991). It allows in principle a complete characterization for the regularization effect of SGD around such manifold -- i.e., the "implicit bias" -- using a stochastic differential equation (SDE) describing the limiting dynamics of the parameters, which is determined jointly by the loss function and the noise covariance. This yields some new results: (1) a global analysis of the implicit bias valid for $η^{-2}$ steps, in contrast to the local analysis of Blanc et al. (2020) that is only valid for $η^{-1.6}$ steps and (2) allowing arbitrary noise covariance. As an application, we show with arbitrary large initialization, label noise SGD can always escape the kernel regime and only requires $O(κ\ln d)$ samples for learning an $κ$-sparse overparametrized linear model in $\mathbb{R}^d$ (Woodworth et al., 2020), while GD initialized in the kernel regime requires $Ω(d)$ samples. This upper bound is minimax optimal and improves the previous $\tilde{O}(κ^2)$ upper bound (HaoChen et al., 2020).

Three FactorsInfluencing Minima in…Three Factors Influencing Minima in SGDThe Anisotropic Noise inStochastic Gradient…The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization EffectsNeural Networks withFinite Intrinsic…Neural Networks with Finite Intrinsic Dimension have no Spurious ValleysTowards Explaining theRegularization Effect o…Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural NetworksGradient DescentProvably Optimizes…Gradient Descent Provably Optimizes Over-parameterized Neural NetworksImplicit Regularizationfor Optimal Sparse…Implicit Regularization for Optimal Sparse RecoveryGradient DescentMaximizes the Margin of…Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksShape Matters:Understanding the…Shape Matters: Understanding the Implicit Bias of the Noise CovarianceLabel Noise SGD ProvablyPrefers Flat Global…Label Noise SGD Provably Prefers Flat Global MinimizersOn the Validity ofModeling SGD with…On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)Implicit Bias of SGD forDiagonal Linear…Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityA Diffusion Theory forDeep Learning Dynamics…A Diffusion Theory for Deep Learning Dynamics: Stochastic Gradient Descent Escapes From Sharp Minima Exponentially FastAnalyzing Sharpnessalong GD Trajectory…Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of StabilityThe Multiscale Structureof Neural Network Loss…The Multiscale Structure of Neural Network Loss Functions: The Effect on Optimization and OriginOn the Trajectories ofSGD Without ReplacementOn the Trajectories of SGD Without ReplacementThe ImplicitRegularization of…The Implicit Regularization of Dynamical Stability in Stochastic Gradient DescentOn the Implicit Bias inDeep-Learning AlgorithmsOn the Implicit Bias in Deep-Learning AlgorithmsA Modern Look at theRelationship between…A Modern Look at the Relationship between Sharpness and GeneralizationUnderstandingMulti-phase Optimizatio…Understanding Multi-phase Optimization Dynamics and Rich Nonlinear Behaviors of ReLU NetworksEdge of StochasticStability: Revisiting…Edge of Stochastic Stability: Revisiting the Edge of Stability for SGDWhy Do You Grok? ATheoretical Analysis of…Why Do You Grok? A Theoretical Analysis of Grokking Modular AdditionHow Neural NetworksLearn the Support is an…How Neural Networks Learn the Support is an Implicit Regularization Effect of SGDUnderstanding theGeneralization Benefits…Understanding the Generalization Benefits of Late Learning Rate DecayWhy Do We Need WeightDecay in Modern Deep…Why Do We Need Weight Decay in Modern Deep Learning?What Happens after SGDReaches Zero Loss? -A…What Happens after SGD Reaches Zero Loss? -A Mathematical Framework過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。