CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

The recently developed vision transformer (ViT) has achieved promising results on image classification compared to convolutional neural networks. Inspired by this, in this paper, we study how to learn multi-scale feature representations in transformer models for image classification. To this end, we propose a dual-branch transformer to combine image patches (i.e., tokens in a transformer) of different sizes to produce stronger image features. Our approach processes small-patch and large-patch tokens with two separate branches of different computational complexity and these tokens are then fused purely by attention multiple times to complement each other. Furthermore, to reduce computation, we develop a simple yet effective token fusion module based on cross attention, which uses a single token for each branch as a query to exchange information with other branches. Our proposed cross-attention only requires linear time for both computational and memory complexity instead of quadratic time otherwise. Extensive experiments demonstrate that our approach performs better than or on par with several concurrent works on vision transformer, in addition to efficient CNN models. For example, on the ImageNet1K dataset, with some architectural changes, our approach outperforms the recent DeiT by a large margin of 2\% with a small to moderate increase in FLOPs and model parameters. Our source codes and models are available at \url{https://github.com/IBM/CrossViT}.

Deep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionFeature Pyramid Networksfor Object DetectionFeature Pyramid Networks for Object DetectionSqueeze-and-ExcitationNetworksSqueeze-and-Excitation NetworksCBAM: ConvolutionalBlock Attention ModuleCBAM: Convolutional Block Attention ModuleQualitative SpatialReasoning over Question…Qualitative Spatial Reasoning over Questions (Short Paper)Selective KernelNetworksSelective Kernel NetworksBig-Little Net: AnEfficient Multi-Scale…Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech RecognitionAttention AugmentedConvolutional NetworksAttention Augmented Convolutional NetworksPyramid VisionTransformer: A Versatil…Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsTransformer inTransformerTransformer in TransformerTokens-to-Token ViT:Training Vision…Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleDual-stream Network forVisual RecognitionDual-stream Network for Visual RecognitionTransGAN: Two PureTransformers Can Make…TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpUFO-ViT: HighPerformance Linear…UFO-ViT: High Performance Linear Vision Transformer without SoftmaxDo You Even NeedAttention? A Stack of…Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNetWave-ViT: UnifyingWavelet and Transformer…Wave-ViT: Unifying Wavelet and Transformers for Visual Representation LearningSPViT: Enabling FasterVision Transformers via…SPViT: Enabling Faster Vision Transformers via Latency-Aware Soft Token PruningA Survey on VisionTransformerA Survey on Vision TransformerViTAEv2: VisionTransformer Advanced by…ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and BeyondRecent progress intransformer-based…Recent progress in transformer-based medical image analysisUnderstandingSelf-attention Mechanis…Understanding Self-attention Mechanism via Dynamical System PerspectiveFcaformer: Forward CrossAttention in Hybrid…Fcaformer: Forward Cross Attention in Hybrid Vision TransformerVision Transformers forImage Classification: A…Vision Transformers for Image Classification: A Comparative SurveyCrossViT:Cross-Attention…CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.