paper-with-me

홈 › Papers

Signature filtering: a lightweight enhancement for statistical watermark detection in large language models

2026-06-16 · Chih-Duo Hong, Yen-Pang Chen, Fang Yu arxiv

Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited. We propose signature filtering, a detection-time module that enhances watermark detection without modifying watermark embedding and text generation. It learns a small set of ``signature'' tokens whose presence makes watermark tests unreliable, and removes these tokens before detection. The signatures are obtained by solving a mixed-integer linear program on a small training set, with constraints that maximize the true positive rate. We additionally derive finite-sample and asymptotic bounds under several attacker models (color-blind, color-adaptive, and distributionally correlated). On four well-known watermark families (Kgw, Sweet, Unigram, Exp), four benchmark corpora (C4, MBPP, HumanEval, Code-Search-Net), and six LLMs (Opt-1.3b, Opt-6.7b, Llama2-13b, Llama3.1-8b, Qwen2.5-14b, Phi-3-medium-14b), 2- and 3-gram signatures raise detection rates in weak-signal and low-entropy settings from 8~31% without filtering to 78~99% with filtering, while keeping false positives controllable and often negligible. In stress tests where we scramble sentences and perturb 25~50% of tokens by dilution, deletions, and substitutions, 2-gram filters for Kgw-style watermarks preserve most of the clean-text detection gains, often matching or outperforming the advanced WinMax watermark detector. Signature filtering thus provides a simple, scalable, and model-agnostic add-on to strengthen watermark-based provenance checks for LLM text in information processing workflows.

📄 PDF Abstract BibTeX arXiv:2606.18430

Code (0)

등록된 구현이 없습니다.

Tasks

Text GenerationText Detection

Similar Papers 제목 키워드 기반

README: Robust Error-Aware Digital Signature Framework via Deep Watermarking Model

2025-07-06 · Hyunwook Choi, Sangyun Won, Daeyeon Hwang, Junhyeok Choi

Deep learning-based watermarking has emerged as a promising solution for robust image authentication and protection. However, existing models are limited by low embedding capacity and vulnerability to bit-level errors, m…

The Stable Signature: Rooting Watermarks in Latent Diffusion Models

2023-03-27 · ICCV 2023 1 · Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze 외

Generative image modeling enables a wide range of applications but raises ethical concerns about responsible deployment. This paper introduces an active strategy combining image watermarking and Latent Diffusion Models. …

Decoder

Stable Signature is Unstable: Removing Image Watermark from Diffusion Models

2024-05-12 · Yuepeng Hu, Zhengyuan Jiang, Moyang Guo, Neil Gong

Watermark has been widely deployed by industry to detect AI-generated images. A recent watermarking framework called \emph{Stable Signature} (proposed by Meta) roots watermark into the parameters of a diffusion model's d…

Decoder

Video Signature: In-generation Watermarking for Latent Video Diffusion Models

2025-05-31 · Yu Huang, JunHao Chen, Qi Zheng, Hanqian Li 외

The rapid development of Artificial Intelligence Generated Content (AIGC) has led to significant progress in video generation but also raises serious concerns about intellectual property protection and reliable content t…

DecoderVideo Generation

Watermarking and Anomaly Detection in Machine Learning Models for LORA RF Fingerprinting

2025-09-18 · Aarushi Mahajan, Wayne Burleson arxiv

Radio frequency fingerprint identification (RFFI) distinguishes wireless devices by the small variations in their analog circuits, avoiding heavy cryptographic authentication. While deep learning on spectrograms improves…

Anomaly Detection