Undetectable Watermarks for Language Models

Recent advances in the capabilities of large language models such as GPT-4 have spurred increasing concern about our ability to detect AI-generated text. Prior works have suggested methods of embedding watermarks in model outputs, by noticeably altering the output distribution. We ask: Is it possible to introduce a watermark without incurring any detectable change to the output distribution? To this end we introduce a cryptographically-inspired notion of undetectable watermarks for language models. That is, watermarks can be detected only with the knowledge of a secret key; without the secret key, it is computationally intractable to distinguish watermarked outputs from those of the original model. In particular, it is impossible for a user to observe any degradation in the quality of the text. Crucially, watermarks should remain undetectable even when the user is allowed to adaptively query the model with arbitrarily chosen prompts. We construct undetectable watermarks based on the existence of one-way functions, a standard assumption in cryptography.

GLTR: StatisticalDetection and…GLTR: Statistical Detection and Visualization of Generated TextAutomatic Detection ofMachine Generated Text…Automatic Detection of Machine Generated Text: A Critical SurveyOn the Possibilities ofAI-Generated Text…On the Possibilities of AI-Generated Text DetectionParaphrasing evadesdetectors of…Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseRobust Multi-bit NaturalLanguage Watermarking…Robust Multi-bit Natural Language Watermarking through Invariant FeaturesGPT detectors are biasedagainst non-native…GPT detectors are biased against non-native English writersDeepTextMark: A DeepLearning-Driven Text…DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated TextPublicly DetectableWatermarking for…Publicly Detectable Watermarking for Language ModelsProvable RobustWatermarking for…Provable Robust Watermarking for AI-Generated TextRobust Distortion-freeWatermarks for Language…Robust Distortion-free Watermarks for Language ModelsA Semantic InvariantRobust Watermark for…A Semantic Invariant Robust Watermark for Large Language ModelsA Resilient andAccessible…A Resilient and Accessible Distribution-Preserving Watermark for Large Language ModelsAdaptive Text Watermarkfor Large Language…Adaptive Text Watermark for Large Language ModelsOn the Learnability ofWatermarks for Language…On the Learnability of Watermarks for Language ModelsAn Unforgeable PubliclyVerifiable Watermark fo…An Unforgeable Publicly Verifiable Watermark for Large Language ModelsUnbiased Watermark forLarge Language ModelsUnbiased Watermark for Large Language ModelsWho Wrote this Code?Watermarking for Code…Who Wrote this Code? Watermarking for Code GenerationAdvancing BeyondIdentification…Advancing Beyond Identification: Multi-bit Watermark for Large Language ModelsIn-Context Watermarksfor Large Language…In-Context Watermarks for Large Language ModelsUndetectable Watermarksfor Language ModelsUndetectable Watermarks for Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。