paper-with-me

홈 › Papers

Revisiting the Robustness of Watermarking to Paraphrasing Attacks

2024-11-08 · Saksham Rastogi, Danish Pruthi

Amidst rising concerns about the internet being proliferated with content generated from language models (LMs), watermarking is seen as a principled way to certify whether text was generated from a model. Many recent watermarking techniques slightly modify the output probabilities of LMs to embed a signal in the generated output that can later be detected. Since early proposals for text watermarking, questions about their robustness to paraphrasing have been prominently discussed. Lately, some techniques are deliberately designed and claimed to be robust to paraphrasing. However, such watermarking schemes do not adequately account for the ease with which they can be reverse-engineered. We show that with access to only a limited number of generations from a black-box watermarked model, we can drastically increase the effectiveness of paraphrasing attacks to evade watermark detection, thereby rendering the watermark ineffective.

📄 PDF Abstract BibTeX arXiv:2411.05277

Code (1)

codeboy5/revisiting-watermark-robustness 공식 구현 pytorch

Similar Papers 제목 키워드 기반

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

2026-05-28 · Yuexin Li, Wenjie Qu, Linyu Wu, Yulin Chen 외 arxiv

Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-based designs remain vulnerable to structural perturbations, such as se…

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks

2026-05-09 · Zhenxin Ai, Haiyun He arxiv

Watermarking for large language models (LLMs) is a promising approach for detecting LLM-generated text and enabling responsible deployment. However, existing watermarking methods are often vulnerable to semantic-invarian…

SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness

2026-05-25 · Jiahao Huo, Wenjie Qu, Yibo Yan, Kening Zheng 외 arxiv

Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globall…

Watermarks for Embeddings-as-a-Service Large Language Models

2025-11-28 · Anudeex Shetty arxiv

Large Language Models (LLMs) have demonstrated exceptional capabilities in natural language understanding and generation. Based on these LLMs, businesses have started to provide Embeddings-as-a-Service (EaaS), offering f…

Natural Language Understanding

PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints

2025-09-25 · Jiahao Huo, Shuliang Liu, Bin Wang, Junyan Zhang 외 arxiv

Semantic-level watermarking (SWM) for large language models (LLMs) enhances watermarking robustness against text modifications and paraphrasing attacks by treating the sentence as the fundamental unit. However, existing …