A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models. A challenge in the domain lies in preserving the distribution of original generated content after watermarking. Our research extends and improves upon existing watermarking framework, placing emphasis on the importance of a \textbf{Di}stribution-\textbf{P}reserving (DiP) watermark. Contrary to the current strategies, our proposed DiPmark simultaneously preserves the original token distribution during watermarking (distribution-preserving), is detectable without access to the language model API and prompts (accessible), and is provably robust to moderate changes of tokens (resilient). DiPmark operates by selecting a random set of tokens prior to the generation of a word, then modifying the token distribution through a distribution-preserving reweight function to enhance the probability of these selected tokens during the sampling process. Extensive empirical evaluation on various language models and tasks demonstrates our approach's distribution-preserving property, accessibility, and resilience, making it a effective solution for watermarking tasks that demand impeccable quality preservation.

BERTScore: EvaluatingText Generation with…BERTScore: Evaluating Text Generation with BERTUndetectable Watermarksfor Language ModelsUndetectable Watermarks for Language ModelsRobust Multi-bit NaturalLanguage Watermarking…Robust Multi-bit Natural Language Watermarking through Invariant FeaturesA Watermark for LargeLanguage ModelsA Watermark for Large Language ModelsParaphrasing evadesdetectors of…Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseOn the Possibilities ofAI-Generated Text…On the Possibilities of AI-Generated Text DetectionLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsDetectGPT: Zero-ShotMachine-Generated Text…DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureUnbiased Watermark forLarge Language ModelsUnbiased Watermark for Large Language ModelsProvable RobustWatermarking for…Provable Robust Watermarking for AI-Generated TextRobust Distortion-freeWatermarks for Language…Robust Distortion-free Watermarks for Language ModelsDeepTextMark: A DeepLearning-Driven Text…DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated TextMark My Words: Analyzingand Evaluating Language…Mark My Words: Analyzing and Evaluating Language Model WatermarksAdaptive Text Watermarkfor Large Language…Adaptive Text Watermark for Large Language ModelsDistortion-freeWatermarks are not Trul…Distortion-free Watermarks are not Truly Distortion-free under Watermark Key CollisionsOn the Learnability ofWatermarks for Language…On the Learnability of Watermarks for Language ModelsOptimizing Watermarksfor Large Language…Optimizing Watermarks for Large Language ModelsMarkLLM: An Open-SourceToolkit for LLM…MarkLLM: An Open-Source Toolkit for LLM WatermarkingWatermarking MakesLanguage Models…Watermarking Makes Language Models RadioactiveA Statistical Frameworkof Watermarks for Large…A Statistical Framework of Watermarks for Large Language Models: Pivot, Detection Efficiency and Optimal RulesDe-mark: WatermarkRemoval in Large…De-mark: Watermark Removal in Large Language ModelsImproved UnbiasedWatermark for Large…Improved Unbiased Watermark for Large Language ModelsTheoretically GroundedFramework for LLM…Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive ApproachAn Ensemble Frameworkfor Unbiased Language…An Ensemble Framework for Unbiased Language Model WatermarkingA Resilient andAccessible…A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。