Pre-Train Your Loss: Easy Bayesian Transfer Learning with Informative Priors

Deep learning is increasingly moving towards a transfer learning paradigm whereby large foundation models are fine-tuned on downstream tasks, starting from an initialization learned on the source task. But an initialization contains relatively little information about the source task. Instead, we show that we can learn highly informative posteriors from the source task, through supervised or self-supervised approaches, which then serve as the basis for priors that modify the whole loss surface on the downstream task. This simple modular approach enables significant performance gains and more data-efficient learning on a variety of downstream classification and segmentation tasks, serving as a drop-in replacement for standard pre-training strategies. These highly informative priors also can be saved for future use, similar to pre-trained weights, and stand in contrast to the zero-mean isotropic uninformative priors that are typically used in Bayesian deep learning.

Rethinking AtrousConvolution for Semanti…Rethinking Atrous Convolution for Semantic Image SegmentationVisualizing the LossLandscape of Neural NetsVisualizing the Loss Landscape of Neural NetsAveraging Weights Leadsto Wider Optima and…Averaging Weights Leads to Wider Optima and Better GeneralizationBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingCyclical StochasticGradient MCMC for…Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningBayesian Deep Learningand a Probabilistic…Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleCoAtNet: MarryingConvolution and…CoAtNet: Marrying Convolution and Attention for All Data SizesWhat Are Bayesian NeuralNetwork Posteriors…What Are Bayesian Neural Network Posteriors Really Like?Dangers of BayesianModel Averaging under…Dangers of Bayesian Model Averaging under Covariate ShiftOn the Opportunities andRisks of Foundation…On the Opportunities and Risks of Foundation ModelsBayesian Neural NetworkPriors RevisitedBayesian Neural Network Priors RevisitedFew-shot Fine-tuning vs.In-context Learning: A…Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and EvaluationAn Information-TheoreticPerspective on…An Information-Theoretic Perspective on Variance-Invariance-Covariance RegularizationBayesian Multi-TaskTransfer Learning for…Bayesian Multi-Task Transfer Learning for Soft Prompt TuningLearning ExpressivePriors for…Learning Expressive Priors for Generalization and Uncertainty Estimation in Neural NetworksThe InterpreterUnderstands Your…The Interpreter Understands Your Meaning: End-to-end Spoken Language Understanding Aided by Speech TranslationCross-Domain TranslationLearning Method…Cross-Domain Translation Learning Method Utilizing Autoencoder Pre-Training for Super-Resolution of Radar Sparse Sensor ArraysTo Compress or Not toCompress -…To Compress or Not to Compress - Self-Supervised Learning and Information Theory: A ReviewIncorporating UnlabelledData into Bayesian…Incorporating Unlabelled Data into Bayesian Neural NetworksTransferring KnowledgeFrom Large Foundation…Transferring Knowledge From Large Foundation Models to Small Downstream ModelsOptimizing postprandialglucose prediction…Optimizing postprandial glucose prediction through integration of diet and exercise: Leveraging transfer learning with imbalanced patient dataFast and Painless ImageReconstruction in Deep…Fast and Painless Image Reconstruction in Deep Image Prior SubspacesFlat Posterior DoesMatter For Bayesian…Flat Posterior Does Matter For Bayesian Transfer LearningPre-Train Your Loss:Easy Bayesian Transfer…Pre-Train Your Loss: Easy Bayesian Transfer Learning with Informative Priors過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。