Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data

The proliferation of generative models, combined with pretraining on web-scale data, raises a timely question: what happens when these models are trained on their own generated outputs? Recent investigations into model-data feedback loops proposed that such loops would lead to a phenomenon termed model collapse, under which performance progressively degrades with each model-data feedback iteration until fitted models become useless. However, those studies largely assumed that new data replace old data over time, where an arguably more realistic assumption is that data accumulate over time. In this paper, we ask: what effect does accumulating data have on model collapse? We empirically study this question by pretraining sequences of language models on text corpora. We confirm that replacing the original real data by each generation's synthetic data does indeed tend towards model collapse, then demonstrate that accumulating the successive generations of synthetic data alongside the original real data avoids model collapse; these results hold across a range of model sizes, architectures, and hyperparameters. We obtain similar results for deep generative models on other types of real data: diffusion models for molecule conformation generation and variational autoencoders for image generation. To understand why accumulating data can avoid model collapse, we use an analytically tractable framework introduced by prior work in which a sequence of linear models are fit to the previous models' outputs. Previous work used this framework to show that if data are replaced, the test error increases with the number of model-fitting iterations; we extend this argument to prove that if data instead accumulate, the test error has a finite upper bound independent of the number of iterations, meaning model collapse no longer occurs.

Fast Kd-Trees for theKullback-Leibler…Fast Kd-Trees for the Kullback-Leibler Divergence and Other Decomposable Bregman DivergencesCombining GenerativeArtificial Intelligence…Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?Large Language ModelsSuffer From Their Own…Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training LoopThe Curse of Recursion:Training on Generated…The Curse of Recursion: Training on Generated Data Makes Models ForgetLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsGPT-4 Technical ReportGPT-4 Technical ReportLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsModel CollapseDemystified: The Case o…Model Collapse Demystified: The Case of RegressionOn the Stability ofIterative Retraining of…On the Stability of Iterative Retraining of Generative Models on their own DataSelf-ConsumingGenerative Models Go MADSelf-Consuming Generative Models Go MADA Tale of Tails: ModelCollapse as a Change of…A Tale of Tails: Model Collapse as a Change of Scaling LawsTowards Understandingthe Interplay of…Towards Understanding the Interplay of Generative Artificial Intelligence and the InternetSelf-ConsumingGenerative Models with…Self-Consuming Generative Models with Curated Data Provably Optimize Human PreferencesReDiFine: ReusableDiffusion Finetuning fo…ReDiFine: Reusable Diffusion Finetuning for Mitigating Degradation in the Chain of DiffusionA linguistic analysis ofundesirable outcomes in…A linguistic analysis of undesirable outcomes in the era of generative AIAnalyzing and ImprovingModel Collapse in…Analyzing and Improving Model Collapse in Rectified Flow ModelsSelf-Improving DiffusionModels with Synthetic…Self-Improving Diffusion Models with Synthetic DataPosition: Model CollapseDoes Not Mean What You…Position: Model Collapse Does Not Mean What You ThinkHow to Synthesize TextData without Model…How to Synthesize Text Data without Model Collapse?Rate of Model Collapsein Recursive TrainingRate of Model Collapse in Recursive TrainingMind the Gap: Examiningthe Self-Improvement…Mind the Gap: Examining the Self-Improvement Capabilities of Large Language ModelsWhat Matters inLLM-generated Data…What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-TuningLLM Web Dynamics:Tracing Model Collapse…LLM Web Dynamics: Tracing Model Collapse in a Network of LLMsRegurgitative Training:The Value of Real Data…Regurgitative Training: The Value of Real Data in Training Large Language ModelsIs Model CollapseInevitable? Breaking th…Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。