Scalable Pre-training of Large Autoregressive Image Models

This paper introduces AIM, a collection of vision models pre-trained with an autoregressive objective. These models are inspired by their textual counterparts, i.e., Large Language Models (LLMs), and exhibit similar scaling properties. Specifically, we highlight two key findings: (1) the performance of the visual features scale with both the model capacity and the quantity of data, (2) the value of the objective function correlates with the performance of the model on downstream tasks. We illustrate the practical implication of these findings by pre-training a 7 billion parameter AIM on 2 billion images, that achieves 84.0% on ImageNet-1k with a frozen trunk. Interestingly, even at this scale, we observe no sign of saturation in performance, suggesting that AIM potentially represents a new frontier for training large-scale vision models. The pre-training of AIM is similar to the pre-training of LLMs, and does not require any image-specific strategy to stabilize the training at scale.

WaveNet: A GenerativeModel for Raw AudioWaveNet: A Generative Model for Raw AudioPixelCNN++: Improvingthe PixelCNN with…PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other ModificationsUnsupervisedRepresentation Learning…Unsupervised Representation Learning by Predicting Image RotationsFixing Weight DecayRegularization in AdamFixing Weight Decay Regularization in AdamLarge Scale GAN Trainingfor High Fidelity…Large Scale GAN Training for High Fidelity Natural Image SynthesisLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersAre Large-scale DatasetsNecessary for…Are Large-scale Datasets Necessary for Self-Supervised Pre-training?LoRA: Low-RankAdaptation of Large…LoRA: Low-Rank Adaptation of Large Language ModelsLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureThe effectiveness of MAEpre-pretraining for…The effectiveness of MAE pre-pretraining for billion-scale pretrainingData Filtering NetworksData Filtering NetworksAn Image is Worth 32Tokens for…An Image is Worth 32 Tokens for Reconstruction and GenerationSapiens: Foundation forHuman Vision ModelsSapiens: Foundation for Human Vision ModelsUnveiling Encoder-FreeVision-Language ModelsUnveiling Encoder-Free Vision-Language ModelsScalingProprioceptive-Visual…Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersMM1: Methods, Analysis &Insights from Multimoda…MM1: Methods, Analysis & Insights from Multimodal LLM Pre-trainingWhen Do We Not NeedLarger Vision Models?When Do We Not Need Larger Vision Models?MultimodalAutoregressive…Multimodal Autoregressive Pre-training of Large Vision EncodersAn Empirical Study ofAutoregressive…An Empirical Study of Autoregressive Pre-Training from VideosAutoregressive Models inVision: A SurveyAutoregressive Models in Vision: A SurveyScaling Laws for OptimalData MixturesScaling Laws for Optimal Data MixturesSmolVLA: AVision-Language-Action…SmolVLA: A Vision-Language-Action Model for Affordable and Efficient RoboticsAutoregressivePretraining with Mamba…Autoregressive Pretraining with Mamba in VisionScalable Pre-training ofLarge Autoregressive…Scalable Pre-training of Large Autoregressive Image ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.