Mirostat: a Neural Text decoding Algorithm that directly controls perplexity

Neural text decoding is important for generating high-quality texts using language models. To generate high-quality text, popular decoding algorithms like top-k, top-p (nucleus), and temperature-based sampling truncate or distort the unreliable low probability tail of the language model. Though these methods generate high-quality text after parameter tuning, they are ad hoc. Not much is known about the control they provide over the statistics of the output, which is important since recent reports show text quality is highest for a specific range of likelihoods. Here, first we provide a theoretical analysis of perplexity in top-k, top-p, and temperature sampling, finding that cross-entropy behaves approximately linearly as a function of p in top-p sampling whereas it is a nonlinear function of k in top-k sampling, under Zipfian statistics. We use this analysis to design a feedback-based adaptive top-k text decoding algorithm called mirostat that generates text (of any length) with a predetermined value of perplexity, and thereby high-quality text without any tuning. Experiments show that for low values of k and p in top-k and top-p sampling, perplexity drops significantly with generated text length, which is also correlated with excessive repetitions in the text (the boredom trap). On the other hand, for large values of k and p, we find that perplexity increases with generated text length, which is correlated with incoherence in the text (confusion trap). Mirostat avoids both traps: experiments show that cross-entropy has a near-linear relation with repetition in generated text. This relation is almost independent of the sampling method but slightly dependent on the model used. Hence, for a given language model, control over perplexity also gives control over repetitions. Experiments with human raters for fluency, coherence, and quality further verify our findings.

The Psycho-Biology ofLanguage: An…The Psycho-Biology of Language: An Introduction to Dynamic PhilologyHuman behavior and theprinciple of least…Human behavior and the principle of least effortArithmetic CodingArithmetic CodingArithmetic Coding forData CompressionArithmetic Coding for Data CompressionAn Estimate of an UpperBound for the Entropy o…An Estimate of an Upper Bound for the Entropy of EnglishFoundations ofstatistical natural…Foundations of statistical natural language processing10.1162/15324430332253322310.1162/153244303322533223Elements of InformationTheoryElements of Information TheoryZipf’s word frequencylaw in natural language…Zipf’s word frequency law in natural language: A critical review and future directionsHierarchical NeuralStory GenerationHierarchical Neural Story GenerationCTRL: A ConditionalTransformer Language…CTRL: A Conditional Transformer Language Model for Controllable GenerationLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersSymbolic MusicGeneration with…Symbolic Music Generation with Transformer-GANsHigh probability or lowinformation? The…High probability or low information? The probability-quality paradox in language generationGrounded Decoding:Guiding Text Generation…Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsArithmetic Sampling:Parallel Diverse…Arithmetic Sampling: Parallel Diverse Decoding for Large Language ModelsDo Language ModelsPlagiarize?Do Language Models Plagiarize?Neural MachineTranslation for Code…Neural Machine Translation for Code GenerationTop-nσ: Not All LogitsAre You NeedTop-nσ: Not All Logits Are You NeedVideoCoT: A VideoChain-of-Thought Datase…VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation ToolScalable Best-of-NSelection for Large…Scalable Best-of-N Selection for Large Language Models via Self-CertaintyVerbalized Sampling: Howto Mitigate Mode…Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM DiversityMeasuring memorizationin language models via…Measuring memorization in language models via probabilistic extractionAdvancing DecodingStrategies: Enhancement…Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMsMirostat: a Neural Textdecoding Algorithm that…Mirostat: a Neural Text decoding Algorithm that directly controls perplexityEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.