Is Space-Time Attention All You Need for Video Understanding?

We present a convolution-free approach to video classification built exclusively on self-attention over space and time. Our method, named "TimeSformer," adapts the standard Transformer architecture to video by enabling spatiotemporal feature learning directly from a sequence of frame-level patches. Our experimental study compares different self-attention schemes and suggests that "divided attention," where temporal attention and spatial attention are separately applied within each block, leads to the best video classification accuracy among the design choices considered. Despite the radically new design, TimeSformer achieves state-of-the-art results on several action recognition benchmarks, including the best reported accuracy on Kinetics-400 and Kinetics-600. Finally, compared to 3D convolutional networks, our model is faster to train, it can achieve dramatically higher test efficiency (at a small drop in accuracy), and it can also be applied to much longer video clips (over one minute long). Code and models are available at: https://github.com/facebookresearch/TimeSformer.

Visualizing Data usingt-SNEVisualizing Data using t-SNEImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseGoing Deeper withConvolutionsGoing Deeper with ConvolutionsAttention Is All YouNeedAttention Is All You NeedRethinkingSpatiotemporal Feature…Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video ClassificationA Short Note aboutKinetics-600A Short Note about Kinetics-600Drop an Octave: ReducingSpatial Redundancy in…Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingHERO: HierarchicalEncoder for…HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleTraining data-efficientimage transformers &…Training data-efficient image transformers & distillation through attentionMultiscale VisionTransformersMultiscale Vision TransformersTransGAN: Two PureTransformers Can Make…TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpPolyViT: Co-trainingVision Transformers on…PolyViT: Co-training Vision Transformers on Images, Videos and AudioEscaping the Big DataParadigm with Compact…Escaping the Big Data Paradigm with Compact TransformersAdaptFormer: AdaptingVision Transformers for…AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionMorphMLP: An EfficientMLP-Like Backbone for…MorphMLP: An Efficient MLP-Like Backbone for Spatial-Temporal Representation LearningEfficient VideoTransformers with…Efficient Video Transformers with Spatial-Temporal Token SelectionWhen Vision TransformersOutperform ResNets…When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsTS2-Net: Token Shift andSelection Transformer…TS2-Net: Token Shift and Selection Transformer for Text-Video RetrievalMPViT: Multi-Path VisionTransformer for Dense…MPViT: Multi-Path Vision Transformer for Dense PredictionDual-AI: Dual-path ActorInteraction Learning fo…Dual-AI: Dual-path Actor Interaction Learning for Group Activity RecognitionVita-CLIP: Video andtext adaptive CLIP via…Vita-CLIP: Video and text adaptive CLIP via Multimodal PromptingIs Space-Time AttentionAll You Need for Video…Is Space-Time Attention All You Need for Video Understanding?Earlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.