Scaling Vision Transformers

Attention-based neural networks such as the Vision Transformer (ViT) have recently attained state-of-the-art results on many computer vision benchmarks. Scale is a primary ingredient in attaining excellent results, therefore, understanding a model's scaling properties is a key to designing future generations effectively. While the laws for scaling Transformer language models have been studied, it is unknown how Vision Transformers scale. To address this, we scale ViT models and data, both up and down, and characterize the relationships between error rate, data, and compute. Along the way, we refine the architecture and training of ViT, reducing memory consumption and increasing accuracy of the resulting models. As a result, we successfully train a ViT model with two billion parameters, which attains a new state-of-the-art on ImageNet of 90.45% top-1 accuracy. The model also performs well for few-shot transfer, for example, reaching 84.86% top-1 accuracy on ImageNet with only 10 examples per class.

Deep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionBig Self-SupervisedModels are Strong…Big Self-Supervised Models are Strong Semi-Supervised LearnersCoAtNet: MarryingConvolution and…CoAtNet: Marrying Convolution and Attention for All Data SizesEfficientNetV2: SmallerModels and Faster…EfficientNetV2: Smaller Models and Faster TrainingDeepViT: Towards DeeperVision TransformerDeepViT: Towards Deeper Vision TransformerTraining data-efficientimage transformers &…Training data-efficient image transformers & distillation through attentionTokens-to-Token ViT:Training Vision…Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetRevisiting ResNets:Improved Training and…Revisiting ResNets: Improved Training and Scaling StrategiesLearning TransferableVisual Models From…Learning Transferable Visual Models From Natural Language SupervisionScaling Up Visual andVision-Language…Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersBottleneck Transformersfor Visual RecognitionBottleneck Transformers for Visual RecognitionCoAtNet: MarryingConvolution and…CoAtNet: Marrying Convolution and Attention for All Data SizesLiT: Zero-Shot Transferwith Locked-image Text…LiT: Zero-Shot Transfer with Locked-image Text TuningSwin Transformer V2:Scaling Up Capacity and…Swin Transformer V2: Scaling Up Capacity and ResolutionTransformers in Vision:A SurveyTransformers in Vision: A SurveyMaxViT: Multi-AxisVision TransformerMaxViT: Multi-Axis Vision TransformerReproducible ScalingLaws for Contrastive…Reproducible Scaling Laws for Contrastive Language-Image LearningViTAEv2: VisionTransformer Advanced by…ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and BeyondAn Overview of VisionTransformers for Image…An Overview of Vision Transformers for Image Processing: A SurveyMasked Autoencoding DoesNot Help Natural…Masked Autoencoding Does Not Help Natural Language Supervision at ScaleVideoMAE V2: ScalingVideo Masked…VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingLearning VisualRepresentations via…Learning Visual Representations via Language-Guided SamplingMultimodal WebNavigation with…Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsScaling VisionTransformersScaling Vision Transformers過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。