paper-with-me

홈 › Papers

Character-Level Perturbations Disrupt LLM Watermarks

2025-09-11 · Zhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, He Zhang, Shirui Pan, Bo Liu, Asif Qumer Gill, Leo Yu Zhang arxiv

Large Language Model (LLM) watermarking embeds detectable signals into generated text for copyright protection, misuse prevention, and content detection. While prior studies evaluate robustness using watermark removal attacks, these methods are often suboptimal, creating the misconception that effective removal requires large perturbations or powerful adversaries. To bridge the gap, we first formalize the system model for LLM watermark, and characterize two realistic threat models constrained on limited access to the watermark detector. We then analyze how different types of perturbation vary in their attack range, i.e., the number of tokens they can affect with a single edit. We observe that character-level perturbations (e.g., typos, swaps, deletions, homoglyphs) can influence multiple tokens simultaneously by disrupting the tokenization process. We demonstrate that character-level perturbations are significantly more effective for watermark removal under the most restrictive threat model. We further propose guided removal attacks based on the Genetic Algorithm (GA) that uses a reference detector for optimization. Under a practical threat model with limited black-box queries to the watermark detector, our method demonstrates strong removal performance. Experiments confirm the superiority of character-level perturbations and the effectiveness of the GA in removing watermarks under realistic constraints. Additionally, we argue there is an adversarial dilemma when considering potential defenses: any fixed defense can be bypassed by a suitable perturbation strategy. Motivated by this principle, we propose an adaptive compound character-level attack. Experimental results show that this approach can effectively defeat the defenses. Our findings highlight significant vulnerabilities in existing LLM watermark schemes and underline the urgency for the development of new robust mechanisms.

📄 PDF Abstract BibTeX arXiv:2509.09112

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FAWA: Fast Adversarial Watermark Attack on Optical Character Recognition (OCR) Systems

2020-12-15 · Lu Chen, Jiao Sun, Wei Xu

Deep neural networks (DNNs) significantly improved the accuracy of optical character recognition (OCR) and inspired many important applications. Unfortunately, OCRs also inherit the vulnerabilities of DNNs under adversar…

Optical Character RecognitionOptical Character Recognition (OCR)

Saliency-Aware Diffusion Reconstruction for Effective Invisible Watermark Removal

2025-04-17 · Inzamamul Alam, MD Tanvir Islam, Simon S. Woo

As digital content becomes increasingly ubiquitous, the need for robust watermark removal techniques has grown due to the inadequacy of existing embedding techniques, which lack robustness. This paper introduces a novel …

Image Restoration

How does Watermarking Affect Visual Language Models in Document Understanding?

2025-04-01 · Chunxue Xu, Yiwei Wang, Bryan Hooi, Yujun Cai 외

Visual Language Models (VLMs) have become foundational models for document understanding tasks, widely used in the processing of complex multimodal documents across domains such as finance, law, and academia. However, do…

document understanding

Evaluating Durability: Benchmark Insights into Multimodal Watermarking

2024-06-06 · JieLin Qiu, William Han, Xuandong Zhao, Shangbang Long 외

With the development of large models, watermarks are increasingly employed to assert copyright, verify authenticity, or monitor content distribution. As applications become more multimodal, the utility of watermarking te…

Text Generation

Disruptive Autoencoders: Leveraging Low-level features for 3D Medical Image Pre-training

2023-07-31 · Jeya Maria Jose Valanarasu, Yucheng Tang, Dong Yang, Ziyue Xu 외

Harnessing the power of pre-training on large-scale datasets like ImageNet forms a fundamental building block for the progress of representation learning-driven solutions in computer vision. Medical images are inherently…

Organ SegmentationRepresentation Learning