VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets. In this paper, we show that video masked autoencoders (VideoMAE) are data-efficient learners for self-supervised video pre-training (SSVP). We are inspired by the recent ImageMAE and propose customized video tube masking with an extremely high ratio. This simple design makes video reconstruction a more challenging self-supervision task, thus encouraging extracting more effective video representations during this pre-training process. We obtain three important findings on SSVP: (1) An extremely high proportion of masking ratio (i.e., 90% to 95%) still yields favorable performance of VideoMAE. The temporally redundant video content enables a higher masking ratio than that of images. (2) VideoMAE achieves impressive results on very small datasets (i.e., around 3k-4k videos) without using any extra data. (3) VideoMAE shows that data quality is more important than data quantity for SSVP. Domain shift between pre-training and target datasets is an important issue. Notably, our VideoMAE with the vanilla ViT can achieve 87.4% on Kinetics-400, 75.4% on Something-Something V2, 91.3% on UCF101, and 62.6% on HMDB51, without using any extra data. Code is available at https://github.com/MCG-NJU/VideoMAE.

UCF101: A Dataset of 101Human Actions Classes…UCF101: A Dataset of 101 Human Actions Classes From Videos in The WildSpatio-temporal videoautoencoder with…Spatio-temporal video autoencoder with differentiable memoryThe Kinetics HumanAction Video DatasetThe Kinetics Human Action Video DatasetSGDR: StochasticGradient Descent with…SGDR: Stochastic Gradient Descent with Warm Restartsmixup: Beyond EmpiricalRisk Minimizationmixup: Beyond Empirical Risk MinimizationLearning SpatiotemporalFeatures via Video and…Learning Spatiotemporal Features via Video and Text Pair DiscriminationVideo RepresentationLearning with Visual…Video Representation Learning with Visual Tempo ConsistencyVIMPAC: VideoPre-Training via Masked…VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive LearningDeepViT: Towards DeeperVision TransformerDeepViT: Towards Deeper Vision TransformerVideoGPT: VideoGeneration using VQ-VAE…VideoGPT: Video Generation using VQ-VAE and TransformersPeCo: PerceptualCodebook for BERT…PeCo: Perceptual Codebook for BERT Pre-training of Vision TransformersAdaptFormer: AdaptingVision Transformers for…AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionUniform Masking:Enabling MAE…Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with LocalityIt Takes Two: MaskedAppearance-Motion…It Takes Two: Masked Appearance-Motion Modeling for Self-supervised Video Transformer Pre-trainingVideoMAE V2: ScalingVideo Masked…VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingMasked Autoencoders inComputer Vision: A…Masked Autoencoders in Computer Vision: A Comprehensive SurveyA Cookbook ofSelf-Supervised LearningA Cookbook of Self-Supervised LearningSTAR-Transformer: ASpatio-temporal Cross…STAR-Transformer: A Spatio-temporal Cross Attention Transformer for Human Action RecognitionVideoLLM: Modeling VideoSequence with Large…VideoLLM: Modeling Video Sequence with Large Language ModelsExploringParameter-Efficient…Exploring Parameter-Efficient Fine-tuning for Improving Communication Efficiency in Federated LearningVisual TuningVisual TuningMamba-360: Survey ofState Space Models as…Mamba-360: Survey of State Space Models as Transformer Alternative for Long Sequence Modelling: Methods, Applications, and ChallengesM-BEV: Masked BEVPerception for Robust…M-BEV: Masked BEV Perception for Robust Autonomous DrivingVideoMAE: MaskedAutoencoders are…VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。