paper-with-me

홈 › Papers

Watermark Smoothing Attacks against Language Models

2024-07-19 · Hongyan Chang, Hamed Hassani, Reza Shokri

Watermarking is a technique used to embed a hidden signal in the probability distribution of text generated by large language models (LLMs), enabling attribution of the text to the originating model. We introduce smoothing attacks and show that existing watermarking methods are not robust against minor modifications of text. An adversary can use weaker language models to smooth out the distribution perturbations caused by watermarks without significantly compromising the quality of the generated text. The modified text resulting from the smoothing attack remains close to the distribution of text that the original model (without watermark) would have produced. Our attack reveals a fundamental limitation of a wide range of watermarking techniques.

📄 PDF Abstract BibTeX arXiv:2407.14206

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Certifiably Robust Image Watermark

2024-07-04 · Zhengyuan Jiang, Moyang Guo, Yuepeng Hu, Jinyuan Jia 외

Generative AI raises many societal concerns such as boosting disinformation and propaganda campaigns. Watermarking AI-generated content is a key technology to address these concerns and has been widely deployed in indust…

Optimizing Adaptive Attacks against Watermarks for Language Models

2024-10-03 · Abdulrahman Diaa, Toluwani Aremu, Nils Lukas

Large Language Models (LLMs) can be misused to spread unwanted content at scale. Content watermarking deters misuse by hiding messages in content, enabling its detection using a secret watermarking key. Robustness is a c…

Misinformation

DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing

2025-10-17 · Ting Qiao, Xing Liu, Wenke Huang, Jianbin Li 외 arxiv

Large web-scale datasets have driven the rapid advancement of pre-trained language models (PLMs), but unauthorized data usage has raised serious copyright concerns. Existing dataset ownership verification (DOV) methods t…

DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack

2025-12-18 · Hao Li, Yubing Ren, Yanan Cao, Yingjie Li 외 arxiv

With the rapid development of cloud-based services, large language models have become increasingly accessible through various web platforms. However, this accessibility has also led to growing risks of model abuse. LLM w…

SoK: How Robust is Image Classification Deep Neural Network Watermarking? (Extended Version)

2021-08-11 · Nils Lukas, Edward Jiang, Xinda Li, Florian Kerschbaum

Deep Neural Network (DNN) watermarking is a method for provenance verification of DNN models. Watermarking should be robust against watermark removal attacks that derive a surrogate model that evades provenance verificat…

image-classificationImage Classification