AI models collapse when trained on recursively generated data

) demonstrated high performance across a variety of language tasks. ChatGPT introduced such language models to the public. It is now clear that generative artificial intelligence (AI) such as large language models (LLMs) is here to stay and will substantially change the ecosystem of online text and images. Here we consider what may happen to GPT-{n} once LLMs contribute much of the text found online. We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear. We refer to this effect as 'model collapse' and show that it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs). We build theoretical intuition behind the phenomenon and portray its ubiquity among all learned generative models. We demonstrate that it must be taken seriously if we are to sustain the benefits of training from large-scale data scraped from the web. Indeed, the value of data collected about genuine human interactions with systems will be increasingly valuable in the presence of LLM-generated content in data crawled from the Internet.

Black Swans and theDomains of StatisticsBlack Swans and the Domains of Statistics5th InternationalConference on Learning…5th International Conference on Learning Representations (ICLR 17)Transformer-BasedFeature Learning for…Transformer-Based Feature Learning for Algorithm Selection in Combinatorial OptimisationThe Curse of Recursion:Training on Generated…The Curse of Recursion: Training on Generated Data Makes Models ForgetScalable watermarkingfor identifying large…Scalable watermarking for identifying large language model outputsHow to intelligentlyembrace generative AI…How to intelligently embrace generative AI: the first guardrails for the use of GenAI in IB researchCan Large ReasoningModels Self-Train?Can Large Reasoning Models Self-Train?Robust Detection ofWatermarks for Large…Robust Detection of Watermarks for Large Language Models Under Human EditsA Survey onLLM-Generated Text…A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future DirectionsHuman-AI agency in theage of generative AIHuman-AI agency in the age of generative AITechno-emotionalprojection in…Techno-emotional projection in human–GenAI relationships: a psychological and ethical conceptual perspectiveAlgorithm Discovery WithLLMs: Evolutionary…Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement LearningRankers, Judges, andAssistants: Towards…Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval EvaluationCapturing a movingtarget: Developing…Capturing a moving target: Developing research on and with AI for Human RelationsAgentStore: ScalableIntegration of…AgentStore: Scalable Integration of Heterogeneous Agents As Specialized Generalist Computer AssistantBARE: Combining Base andInstruction-Tuned…BARE: Combining Base and Instruction-Tuned Language Models for Better Synthetic Data GenerationAI models collapse whentrained on recursively…AI models collapse when trained on recursively generated data過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。