Self-Consuming Generative Models Go MAD

Seismic advances in generative AI algorithms for imagery, text, and other data types has led to the temptation to use synthetic data to train next-generation models. Repeating this process creates an autophagous (self-consuming) loop whose properties are poorly understood. We conduct a thorough analytical and empirical analysis using state-of-the-art generative image models of three families of autophagous loops that differ in how fixed or fresh real training data is available through the generations of training and in whether the samples from previous generation models have been biased to trade off data quality versus diversity. Our primary conclusion across all scenarios is that without enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decrease. We term this condition Model Autophagy Disorder (MAD), making analogy to mad cow disease.

HierarchicalText-Conditional Image…Hierarchical Text-Conditional Image Generation with CLIP LatentsWill we run out of data?An analysis of the…Will we run out of data? An analysis of the limits of scaling datasets in Machine LearningCombining GenerativeArtificial Intelligence…Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?The Curse of Recursion:Training on Generated…The Curse of Recursion: Training on Generated Data Makes Models ForgetA data augmentationperspective on diffusio…A data augmentation perspective on diffusion models and retrievalSelf-Instruct: AligningLanguage Models with…Self-Instruct: Aligning Language Models with Self-Generated InstructionsGPT-4 Technical ReportGPT-4 Technical ReportLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsMusicLM: GeneratingMusic From TextMusicLM: Generating Music From TextLeaving Reality toImagination: Robust…Leaving Reality to Imagination: Robust Classification via Generated DatasetsTowards Understandingthe Interplay of…Towards Understanding the Interplay of Generative Artificial Intelligence and the InternetOn the Reliability ofWatermarks for Large…On the Reliability of Watermarks for Large Language ModelsLarge Language ModelsSuffer From Their Own…Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training LoopNepotistically TrainedGenerative-AI Models…Nepotistically Trained Generative-AI Models CollapseIs Model CollapseInevitable? Breaking th…Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic DataModel CollapseDemystified: The Case o…Model Collapse Demystified: The Case of RegressionA Tale of Tails: ModelCollapse as a Change of…A Tale of Tails: Model Collapse as a Change of Scaling LawsTowards TheoreticalUnderstandings of…Towards Theoretical Understandings of Self-Consuming Generative ModelsSelf-Improving DiffusionModels with Synthetic…Self-Improving Diffusion Models with Synthetic DataReDiFine: ReusableDiffusion Finetuning fo…ReDiFine: Reusable Diffusion Finetuning for Mitigating Degradation in the Chain of DiffusionWhen AI Eats Itself: Onthe Caveats of Data…When AI Eats Itself: On the Caveats of Data Pollution in the Era of Generative AIBeyond Model Collapse:Scaling Up with…Beyond Model Collapse: Scaling Up with Synthesized Data Requires ReinforcementPosition: Model CollapseDoes Not Mean What You…Position: Model Collapse Does Not Mean What You ThinkStrong Model CollapseStrong Model CollapseSelf-ConsumingGenerative Models Go MADSelf-Consuming Generative Models Go MADEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.