Adam: A Method for Stochastic Optimization

We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments. The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters. The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients. The hyper-parameters have intuitive interpretations and typically require little tuning. Some connections to related algorithms, on which Adam was inspired, are discussed. We also analyze the theoretical convergence properties of the algorithm and provide a regret bound on the convergence rate that is comparable to the best known results under the online convex optimization framework. Empirical results demonstrate that Adam works well in practice and compares favorably to other stochastic optimization methods. Finally, we discuss AdaMax, a variant of Adam based on the infinity norm.

Reducing theDimensionality of Data…Reducing the Dimensionality of Data with Neural NetworksAdaptive SubgradientMethods for Online…Adaptive Subgradient Methods for Online Learning and Stochastic OptimizationNon-Asymptotic Analysisof Stochastic…Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Machine LearningLearning Word Vectorsfor Sentiment AnalysisLearning Word Vectors for Sentiment AnalysisDeep Neural Networks forAcoustic Modeling in…Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research GroupsADADELTA: An AdaptiveLearning Rate MethodADADELTA: An Adaptive Learning Rate MethodImproving neuralnetworks by preventing…Improving neural networks by preventing co-adaptation of feature detectorsOn the importance ofinitialization and…On the importance of initialization and momentum in deep learningSpeech Recognition withDeep Recurrent Neural…Speech Recognition with Deep Recurrent Neural NetworksLOL: An Investigationinto Cybernetic Humor…LOL: An Investigation into Cybernetic Humor, or: Can Machines Laugh?Recent advances in deeplearning for speech…Recent advances in deep learning for speech research at MicrosoftNo More Pesky LearningRatesNo More Pesky Learning RatesA Convolutional NeuralNetwork for Automatic…A Convolutional Neural Network for Automatic Characterization of Plaque Composition in Carotid UltrasoundA Genetic ProgrammingApproach to Designing…A Genetic Programming Approach to Designing Convolutional Neural Network ArchitecturesMultitask Learning forCross-Domain Image…Multitask Learning for Cross-Domain Image CaptioningComparable Study OfModeling Units For…Comparable Study Of Modeling Units For End-To-End Mandarin Speech RecognitionSDDNet: Real-Time CrackSegmentationSDDNet: Real-Time Crack SegmentationCrowd Counting andDensity Estimation by…Crowd Counting and Density Estimation by Trellis Encoder-Decoder NetworksDeep Packet: A NovelApproach For Encrypted…Deep Packet: A Novel Approach For Encrypted Traffic Classification Using Deep LearningLow-resource Deep EntityResolution with Transfe…Low-resource Deep Entity Resolution with Transfer and Active LearningDeep Weighted AveragingClassifiersDeep Weighted Averaging ClassifiersDensePhysNet: LearningDense Physical Object…DensePhysNet: Learning Dense Physical Object Representations Via Multi-Step Dynamic InteractionsPFA-GAN: ProgressiveFace Aging With…PFA-GAN: Progressive Face Aging With Generative Adversarial NetworkSupervised Learning withProjected Entangled Pai…Supervised Learning with Projected Entangled Pair StatesAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic Optimization過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。