Video Representation Learning by Dense Predictive Coding

The objective of this paper is self-supervised learning of spatio-temporal embeddings from video, suitable for human action recognition. We make three contributions: First, we introduce the Dense Predictive Coding (DPC) framework for self-supervised representation learning on videos. This learns a dense encoding of spatio-temporal blocks by recurrently predicting future representations; Second, we propose a curriculum training scheme to predict further into the future with progressively less temporal context. This encourages the model to only encode slowly varying spatial-temporal signals, therefore leading to semantic representations; Third, we evaluate the approach by first training the DPC model on the Kinetics-400 dataset with self-supervised learning, and then finetuning the representation on a downstream task, i.e. action recognition. With single stream (RGB only), DPC pretrained representations achieve state-of-the-art self-supervised performance on both UCF101(75.7% top1 acc) and HMDB51(35.7% top1 acc), outperforming all previous learning methods by a significant margin, and approaching the performance of a baseline pre-trained on ImageNet.

DistributedRepresentations of Word…Distributed Representations of Words and Phrases and their CompositionalityUnsupervised Learning ofVideo Representations…Unsupervised Learning of Video Representations using LSTMsUnsupervised Learning ofVisual Representations…Unsupervised Learning of Visual Representations by Solving Jigsaw PuzzlesGenerating Videos withScene DynamicsGenerating Videos with Scene DynamicsSelf-Supervised VideoRepresentation Learning…Self-Supervised Video Representation Learning With Odd-One-Out NetworksThe Kinetics HumanAction Video DatasetThe Kinetics Human Action Video DatasetSelf-supervisedSpatiotemporal Feature…Self-supervised Spatiotemporal Feature Learning by Video Geometric TransformationsRepresentation Learningwith Contrastive…Representation Learning with Contrastive Predictive CodingLearning and Using theArrow of TimeLearning and Using the Arrow of TimeDeep Clustering forUnsupervised Learning o…Deep Clustering for Unsupervised Learning of Visual FeaturesSelf-supervised Learningfor Video Correspondenc…Self-supervised Learning for Video Correspondence FlowSlowFast Networks forVideo RecognitionSlowFast Networks for Video RecognitionContrastiveBidirectional…Contrastive Bidirectional Transformer for Temporal Representation LearningMemory-Augmented DensePredictive Coding for…Memory-Augmented Dense Predictive Coding for Video Representation LearningSelf-Supervised Learningof Video-Induced Visual…Self-Supervised Learning of Video-Induced Visual InvariancesLearning SpatiotemporalFeatures via Video and…Learning Spatiotemporal Features via Video and Text Pair DiscriminationMulti-modalSelf-Supervision from…Multi-modal Self-Supervision from Generalized Data TransformationsEnd-to-End Learning ofVisual Representations…End-to-End Learning of Visual Representations From Uncurated Instructional VideosSelf-Supervised VideoRepresentation Using…Self-Supervised Video Representation Using Pretext-Contrastive LearningVideo Understanding asMachine TranslationVideo Understanding as Machine TranslationContrastive predictivecoding with transformer…Contrastive predictive coding with transformer for video representation learningComposable AugmentationEncoding for Video…Composable Augmentation Encoding for Video Representation LearningEnhancing UnsupervisedVideo Representation…Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the MotionMotion-awareSelf-supervised Video…Motion-aware Self-supervised Video Representation Learning via Foreground-background MergingVideo RepresentationLearning by Dense…Video Representation Learning by Dense Predictive CodingEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.