Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Stochastic neurons and hard non-linearities can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic or non-smooth neurons? I.e., can we "back-propagate" through these stochastic neurons? We examine this question, existing approaches, and compare four families of solutions, applicable in different settings. One of them is the minimum variance unbiased gradient estimator for stochatic binary neurons (a special case of the REINFORCE algorithm). A second approach, introduced here, decomposes the operation of a binary stochastic neuron into a stochastic binary part and a smooth differentiable part, which approximates the expected effect of the pure stochatic binary neuron to first order. A third approach involves the injection of additive or multiplicative noise in a computational graph that is otherwise differentiable. A fourth approach heuristically copies the gradient with respect to the stochastic output directly as an estimator of the gradient with respect to the sigmoid argument (we call this the straight-through estimator). To explore a context where these estimators are useful, we consider a small-scale version of {\em conditional computation}, where sparse stochastic units form a distributed representation of gaters that can turn off in combinatorially many ways large chunks of the computation performed in the rest of the neural network. In this case, it is important that the gating units produce an actual 0 most of the time. The resulting sparsity can be potentially be exploited to greatly reduce the computational cost of large deep networks for which conditional computation would be useful.

Learning Representationsby Back-Propagating…Learning Representations by Back-Propagating ErrorsMultivariate stochasticapproximation using a…Multivariate stochastic approximation using a simultaneous perturbation gradient approximationHierarchical RecurrentNeural Networks for…Hierarchical Recurrent Neural Networks for Long-Term DependenciesThe Optimal RewardBaseline for…The Optimal Reward Baseline for Gradient-Based Reinforcement LearningGradient Learning inSpiking Neural Networks…Gradient Learning in Spiking Neural Networks by Dynamic Perturbation of ConductancesExtracting and composingrobust features with…Extracting and composing robust features with denoising autoencodersSemantic hashingSemantic hashingImproving neuralnetworks by preventing…Improving neural networks by preventing co-adaptation of feature detectorsDeep Learning ofRepresentations: Lookin…Deep Learning of Representations: Looking Forwardopenalex_id:w2949227999openalex_id:w2949227999Emergence of Languagewith Multi-agent Games…Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of SymbolsLearning Sparse NeuralNetworks through L_0…Learning Sparse Neural Networks through L_0 RegularizationRouting Networks and theChallenges of Modular…Routing Networks and the Challenges of Modular and Compositional ComputationGenerative AdversarialNetworks (GANs)…Generative Adversarial Networks (GANs): Challenges, Solutions, and Future DirectionsQ-BERT: Hessian BasedUltra Low Precision…Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTLearning Multi-granularQuantized Embeddings fo…Learning Multi-granular Quantized Embeddings for Large-Vocab Categorical Features in Recommender SystemsA Free Lunch From ANN:Towards Efficient…A Free Lunch From ANN: Towards Efficient, Accurate Spiking Neural Networks CalibrationDegree-Quant:Quantization-Aware…Degree-Quant: Quantization-Aware Training for Graph Neural NetworksDenoising Self-AttentiveSequential…Denoising Self-Attentive Sequential RecommendationMERF: Memory-EfficientRadiance Fields for…MERF: Memory-Efficient Radiance Fields for Real-time View Synthesis in Unbounded ScenesSTORM: EfficientStochastic Transformer…STORM: Efficient Stochastic Transformer based World Models for Reinforcement LearningEstimating orPropagating Gradients…Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。