On the Reliability of Watermarks for Large Language Models

As LLMs become commonplace, machine-generated text has the potential to flood the internet with spam, social media bots, and valueless content. Watermarking is a simple and effective strategy for mitigating such harms by enabling the detection and documentation of LLM-generated text. Yet a crucial question remains: How reliable is watermarking in realistic settings in the wild? There, watermarked text may be modified to suit a user's needs, or entirely rewritten to avoid detection. We study the robustness of watermarked text after it is re-written by humans, paraphrased by a non-watermarked LLM, or mixed into a longer hand-written document. We find that watermarks remain detectable even after human and machine paraphrasing. While these attacks dilute the strength of the watermark, paraphrases are statistically likely to leak n-grams or even longer fragments of the original text, resulting in high-confidence detections when enough tokens are observed. For example, after strong human paraphrasing the watermark is detectable after observing 800 tokens on average, when setting a 1e-5 false positive rate. We also consider a range of new detection schemes that are sensitive to short spans of watermarked text embedded inside a large document, and we compare the robustness of watermarking to other kinds of detectors.

GLTR: StatisticalDetection and…GLTR: Statistical Detection and Visualization of Generated TextA Watermark for LargeLanguage ModelsA Watermark for Large Language ModelsParaphrasing evadesdetectors of…Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseCan AI-Generated Text beReliably Detected?Can AI-Generated Text be Reliably Detected?GPT detectors are biasedagainst non-native…GPT detectors are biased against non-native English writersDetectGPT: Zero-ShotMachine-Generated Text…DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureOn the Possibilities ofAI-Generated Text…On the Possibilities of AI-Generated Text DetectionThree Bricks toConsolidate Watermarks…Three Bricks to Consolidate Watermarks for Large Language ModelsThe Science of DetectingLLM-Generated TextsThe Science of Detecting LLM-Generated TextsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsWho Wrote this Code?Watermarking for Code…Who Wrote this Code? Watermarking for Code GenerationWatermarks in the Sand:Impossibility of Strong…Watermarks in the Sand: Impossibility of Strong Watermarking for Generative ModelsParaphrasing evadesdetectors of…Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseTowards OptimalStatistical WatermarkingTowards Optimal Statistical WatermarkingAdvancing BeyondIdentification…Advancing Beyond Identification: Multi-bit Watermark for Large Language ModelsA Robust Semantics-basedWatermark for Large…A Robust Semantics-based Watermark for Large Language Model against ParaphrasingAdaptive Text Watermarkfor Large Language…Adaptive Text Watermark for Large Language ModelsTowards CodableWatermarking for…Towards Codable Watermarking for Injecting Multi-Bits Information to LLMsOn the Learnability ofWatermarks for Language…On the Learnability of Watermarks for Language ModelsWho Wrote this Code?Watermarking for Code…Who Wrote this Code? Watermarking for Code GenerationWaterBench: TowardsHolistic Evaluation of…WaterBench: Towards Holistic Evaluation of Watermarks for Large Language ModelsToken-SpecificWatermarking with…Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language ModelsWAPITI: A Watermark forFinetuned Open-Source…WAPITI: A Watermark for Finetuned Open-Source LLMsOn the Reliability ofWatermarks for Large…On the Reliability of Watermarks for Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。