Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs

Large Language Models (LLMs) generate text by sampling the next token from a probability distribution over the vocabulary at each decoding step. Popular sampling methods like top-p (nucleus sampling) often struggle to balance quality and diversity, especially at higher temperatures which lead to incoherent or repetitive outputs. We propose min-p sampling, a dynamic truncation method that adjusts the sampling threshold based on the model's confidence by using the top token's probability as a scaling factor. Our experiments on benchmarks including GPQA, GSM8K, and AlpacaEval Creative Writing show that min-p sampling improves both the quality and diversity of generated text across different model families (Mistral and Llama 3) and model sizes (1B to 123B parameters), especially at higher temperatures. Human evaluations further show a clear preference for min-p sampling, in both text quality and creativity. Min-p sampling has been adopted by popular open-source LLM frameworks, including Hugging Face Transformers, VLLM, and many others, highlighting its considerable impact on improving text generation quality.

Beam Search Strategiesfor Neural Machine…Beam Search Strategies for Neural Machine TranslationHierarchical NeuralStory GenerationHierarchical Neural Story GenerationTransformers:State-of-the-Art Natura…Transformers: State-of-the-Art Natural Language ProcessingTruncation Sampling asLanguage Model…Truncation Sampling as Language Model DesmoothingOn Decoding Strategiesfor Neural Text…On Decoding Strategies for Neural Text GeneratorsSelf-ConsistencyImproves Chain of…Self-Consistency Improves Chain of Thought Reasoning in Language ModelsAdaptive Decoding viaLatent Preference…Adaptive Decoding via Latent Preference OptimizationThe Llama 3 Herd ofModelsThe Llama 3 Herd of ModelsTop-nσ: Not All LogitsAre You NeedTop-nσ: Not All Logits Are You NeedImproving Open-EndedText Generation via…Improving Open-Ended Text Generation via Adaptive DecodingChain-of-ThoughtReasoning Without…Chain-of-Thought Reasoning Without PromptingDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningTop-nσ: Not All LogitsAre You NeedTop-nσ: Not All Logits Are You NeedDiverse PreferenceOptimizationDiverse Preference OptimizationVerbalized Sampling: Howto Mitigate Mode…Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM DiversityOptimizing Temperaturefor Language Models wit…Optimizing Temperature for Language Models with Multi-Sample Inferencep-less Sampling: ARobust…p-less Sampling: A Robust Hyperparameter-Free Approach for LLM DecodingSample Smart, Not Hard:Correctness-First…Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMsGeneralizedInterpolating Discrete…Generalized Interpolating Discrete DiffusionBalancing Diversity andRisk in LLM Sampling…Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text GenerationDPWriter: ReinforcementLearning with Diverse…DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative WritingUpSkill: MutualInformation Skill…UpSkill: Mutual Information Skill Learning for Structured Response Diversity in LLMsDecoding as Optimisationon the Probability…Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K SamplersLess is More: ImprovingLLM Reasoning with…Less is More: Improving LLM Reasoning with Minimal Test-Time InterventionTurning Up the Heat:Min-p Sampling for…Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。