Memory-Augmented Dense Predictive Coding for Video Representation Learning

The objective of this paper is self-supervised learning from video, in particular for representations for action recognition. We make the following contributions: (i) We propose a new architecture and learning framework Memory-augmented Dense Predictive Coding (MemDPC) for the task. It is trained with a predictive attention mechanism over the set of compressed memories, such that any future states can always be constructed by a convex combination of the condense representations, allowing to make multiple hypotheses efficiently. (ii) We investigate visual-only self-supervised video representation learning from RGB frames, or from unsupervised optical flow, or both. (iii) We thoroughly evaluate the quality of learnt representation on four different downstream tasks: action recognition, video retrieval, learning with scarce annotations, and unintentional action classification. In all cases, we demonstrate state-of-the-art or comparable performance over other approaches with orders of magnitude fewer training data.

Self-supervisedSpatiotemporal Feature…Self-supervised Spatiotemporal Feature Learning by Video Geometric TransformationsVideo RepresentationLearning by Dense…Video Representation Learning by Dense Predictive CodingContrastiveBidirectional…Contrastive Bidirectional Transformer for Temporal Representation LearningData-Efficient ImageRecognition with…Data-Efficient Image Recognition with Contrastive Predictive CodingSelf-supervised Learningfor Video Correspondenc…Self-supervised Learning for Video Correspondence FlowSelf-SupervisedSpatiotemporal Learning…Self-Supervised Spatiotemporal Learning via Video Clip Order PredictionMulti-modalSelf-Supervision from…Multi-modal Self-Supervision from Generalized Data TransformationsContrastive MultiviewCodingContrastive Multiview CodingVideo Cloze Procedurefor Self-Supervised…Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningSelf-Supervised Learningby Cross-Modal…Self-Supervised Learning by Cross-Modal Audio-Video ClusteringMAST: A Memory-AugmentedSelf-Supervised TrackerMAST: A Memory-Augmented Self-Supervised TrackerMomentum Contrast forUnsupervised Visual…Momentum Contrast for Unsupervised Visual Representation LearningCan Temporal InformationHelp with Contrastive…Can Temporal Information Help with Contrastive Self-Supervised Learning?SpatiotemporalContrastive Video…Spatiotemporal Contrastive Video Representation LearningTime-EquivariantContrastive Video…Time-Equivariant Contrastive Video Representation LearningSelf-Supervised VideoRepresentation Learning…Self-Supervised Video Representation Learning with Meta-Contrastive NetworkRemoving the Backgroundby Adding the…Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation LearningVATT: Transformers forMultimodal…VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextContrastive Learning ofImage Representations…Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyComposable AugmentationEncoding for Video…Composable Augmentation Encoding for Video Representation LearningMotion-awareSelf-supervised Video…Motion-aware Self-supervised Video Representation Learning via Foreground-background MergingTCLR: TemporalContrastive Learning fo…TCLR: Temporal Contrastive Learning for Video RepresentationHierarchically DecoupledSpatial-Temporal…Hierarchically Decoupled Spatial-Temporal Contrast for Self-supervised Video Representation LearningControllableAugmentations for Video…Controllable Augmentations for Video Representation LearningMemory-Augmented DensePredictive Coding for…Memory-Augmented Dense Predictive Coding for Video Representation LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.