The Curious Case of Neural Text Degeneration

Despite considerable advances in neural language modeling, it remains an open question what the best decoding strategy is for text generation from a language model (e.g. to generate a story). The counter-intuitive empirical observation is that even though the use of likelihood as training objective leads to high quality models for a broad range of language understanding tasks, maximization-based decoding methods such as beam search lead to degeneration — output text that is bland, incoherent, or gets stuck in repetitive loops. To address this we propose Nucleus Sampling, a simple but effective method to draw considerably higher quality text out of neural language models. Our approach avoids text degeneration by truncating the unreliable tail of the probability distribution, sampling from the dynamic nucleus of tokens containing the vast majority of the probability mass. To properly examine current maximization-based and stochastic decoding methods, we compare generations from each of these methods to the distribution of human text along several axes such as likelihood, diversity, and repetition. Our results show that (1) maximization is an inappropriate decoding objective for open-ended text generation, (2) the probability distributions of the best current language models have an unreliable tail which needs to be truncated during generation and (3) Nucleus Sampling is the best decoding strategy for generating long-form text that is both high-quality — as measured by human evaluation — and as diverse as human-written text.

openalex_id:w2963506925openalex_id:w2963506925Logic and ConversationLogic and ConversationA Learning Algorithm forBoltzmann MachinesA Learning Algorithm for Boltzmann MachinesLSTM Neural Networks forLanguage ModelingLSTM Neural Networks for Language ModelingOn the Properties ofNeural Machine…On the Properties of Neural Machine Translation: Encoder-Decoder ApproachesEffective Approaches toAttention-based Neural…Effective Approaches to Attention-based Neural Machine TranslationDeep LearningDeep LearningDeep ReinforcementLearning for Dialogue…Deep Reinforcement Learning for Dialogue GenerationChallenges inData-to-Document…Challenges in Data-to-Document GenerationTexygen: A BenchmarkingPlatform for Text…Texygen: A Benchmarking Platform for Text Generation ModelsOn NMT Search Errors andModel Errors: Cat Got…On NMT Search Errors and Model Errors: Cat Got Your Tongue?Language GANs FallingShortLanguage GANs Falling Shortopenalex_id:w3102877762openalex_id:w3102877762Training Language GANsfrom ScratchTraining Language GANs from ScratchGeDi: GenerativeDiscriminator Guided…GeDi: Generative Discriminator Guided Sequence GenerationComputer-Aided Design asLanguageComputer-Aided Design as LanguageCPM: A Large-scaleGenerative Chinese…CPM: A Large-scale Generative Chinese Pre-trained Language ModelEnd-to-End generation ofMultiple-Choice…End-to-End generation of Multiple-Choice questions using Text-to-Text transfer Transformer modelsHigh probability or lowinformation? The…High probability or low information? The probability-quality paradox in language generationDraft, Sketch, andProve: Guiding Formal…Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal ProofsCan Large LanguageModels Be an Alternativ…Can Large Language Models Be an Alternative to Human Evaluations?Evaluating largelanguage models for use…Evaluating large language models for use in healthcare: A framework for translational value assessmentNavigating theComplexity of Generativ…Navigating the Complexity of Generative AI Adoption in Software EngineeringMART: Improving LLMSafety with Multi-round…MART: Improving LLM Safety with Multi-round Automatic Red-TeamingThe Curious Case ofNeural Text DegenerationThe Curious Case of Neural Text DegenerationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.