paper-with-me

홈 › Papers

Watermark under Fire: A Robustness Evaluation of LLM Watermarking

2024-11-20 · Jiacheng Liang, Zian Wang, Lauren Hong, Shouling Ji, Ting Wang

Various watermarking methods (``watermarkers'') have been proposed to identify LLM-generated texts; yet, due to the lack of unified evaluation platforms, many critical questions remain under-explored: i) What are the strengths/limitations of various watermarkers, especially their attack robustness? ii) How do various design choices impact their robustness? iii) How to optimally operate watermarkers in adversarial environments? To fill this gap, we systematize existing LLM watermarkers and watermark removal attacks, mapping out their design spaces. We then develop WaterPark, a unified platform that integrates 10 state-of-the-art watermarkers and 12 representative attacks. More importantly, by leveraging WaterPark, we conduct a comprehensive assessment of existing watermarkers, unveiling the impact of various design choices on their attack robustness. We further explore the best practices to operate watermarkers in adversarial environments. We believe our study sheds light on current LLM watermarking techniques while WaterPark serves as a valuable testbed to facilitate future research.

📄 PDF Abstract BibTeX arXiv:2411.13425

Code (1)

JACKPURCELL/sok-llm-watermark 공식 구현 pytorch

Tasks

Language ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

Evaluating the Robustness of Trigger Set-Based Watermarks Embedded in Deep Neural Networks

2021-06-18 · Suyoung Lee, Wonho Song, Suman Jana, Meeyoung Cha 외

Trigger set-based watermarking schemes have gained emerging attention as they provide a means to prove ownership for deep neural network model owners. In this paper, we argue that state-of-the-art trigger set-based water…

Enhancing Robustness in Post-Processing Watermarking: An Ensemble Attack Network Using CNNs and Transformers

2025-09-03 · Tzuhsuan Huang, Cheng Yu Yeo, Tsai-Ling Huang, Hong-Han Shuai 외 arxiv

Recent studies on deep watermarking have predominantly focused on in-processing watermarking, which integrates the watermarking process into image generation. However, post-processing watermarking, which embeds watermark…

Image Generation

New Evaluation Metrics Capture Quality Degradation due to LLM Watermarking

2023-12-04 · Karanpartap Singh, James Zou

With the increasing use of large-language models (LLMs) like ChatGPT, watermarking has emerged as a promising approach for tracing machine-generated content. However, research on LLM watermarking often relies on simple p…

Binary ClassificationDiversity

Optimizing Adaptive Attacks against Watermarks for Language Models

2024-10-03 · Abdulrahman Diaa, Toluwani Aremu, Nils Lukas

Large Language Models (LLMs) can be misused to spread unwanted content at scale. Content watermarking deters misuse by hiding messages in content, enabling its detection using a secret watermarking key. Robustness is a c…

Misinformation

SoK: How Robust is Image Classification Deep Neural Network Watermarking? (Extended Version)

2021-08-11 · Nils Lukas, Edward Jiang, Xinda Li, Florian Kerschbaum

Deep Neural Network (DNN) watermarking is a method for provenance verification of DNN models. Watermarking should be robust against watermark removal attacks that derive a surrogate model that evades provenance verificat…

image-classificationImage Classification