paper-with-me

홈 › Papers

Black-Box Detection of Language Model Watermarks

2024-05-28 · Thibaud Gloaguen, Nikola Jovanović, Robin Staab, Martin Vechev

Watermarking has emerged as a promising way to detect LLM-generated text, by augmenting LLM generations with later detectable signals. Recent work has proposed multiple families of watermarking schemes, several of which focus on preserving the LLM distribution. This distribution-preservation property is motivated by the fact that it is a tractable proxy for retaining LLM capabilities, as well as the inherently implied undetectability of the watermark by downstream users. Yet, despite much discourse around undetectability, no prior work has investigated the practical detectability of any of the current watermarking schemes in a realistic black-box setting. In this work we tackle this for the first time, developing rigorous statistical tests to detect the presence, and estimate parameters, of all three popular watermarking scheme families, using only a limited number of black-box queries. We experimentally confirm the effectiveness of our methods on a range of schemes and a diverse set of open-source models. Further, we validate the feasibility of our tests on real-world APIs. Our findings indicate that current watermarking schemes are more detectable than previously believed.

📄 PDF Abstract BibTeX arXiv:2405.20777

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited Data

2023-09-07 · Yifan Lu, Wenxuan Li, Mi Zhang, Xudong Pan 외

To protect the intellectual property of well-trained deep neural networks (DNNs), black-box watermarks, which are embedded into the prediction behavior of DNN models on a set of specially-crafted samples and extracted fr…

Proving membership in LLM pretraining data via data watermarks

2024-02-16 · Johnny Tian-Zheng Wei, Ryan Yixiang Wang, Robin Jia

Detecting whether copyright holders' works were used in LLM pretraining is poised to be an important problem. This work proposes using data watermarks to enable principled detection with only black-box model access, prov…

SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery Attacks

2026-01-28 · Xin Zhang, Zijin Yang, Kejiang Chen, Linfeng Ma 외 arxiv

Latent-based watermarks, integrated into the generation process of latent diffusion models (LDMs), simplify detection and attribution of generated images. However, recent black-box forgery attacks, where an attacker need…

Contrastive Learning

Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective

2026-06-29 · Cheng-Yi Lee, Yichi Zhang, Yuchen Yang, Chun-Shien Lu 외 arxiv

Recent studies have shown that semantic watermarks, which embed information into the initial noise of latent diffusion models (LDMs), are vulnerable to black-box forgery attacks. However, existing methods primarily rely …

FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks

2025-04-13 · Tianyi Wang, Harry Cheng, Ming-Hui Liu, Mohan Kankanhalli

Proactive Deepfake detection via robust watermarks has been raised ever since passive Deepfake detectors encountered challenges in identifying high-quality synthetic images. However, while demonstrating reasonable detect…

DeepFake DetectionFace Swapping