paper-with-me

홈 › Papers

Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

2026-05-28 · Zhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen, Runcong Zhao arxiv

Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access multiple models (today's reality), watermarks trivially fail. Watermarks perturb output distributions away from the original, and in competitive markets, these perturbations are typically independent across providers. We theoretically prove that averaging output probability distributions recovers the unwatermarked distribution with up to a second-order error term. Empirically, simply averaging 3-5 models cancels out these perturbations. We introduce WASH (Watermark Attenuation via Statistical Hybridisation), which solves practical challenges in ensemble generation: vocabulary misalignment and tokenisation differences across heterogeneous models. Experiments across six watermarking schemes and three LLMs show that averaging across 3 models suppresses detection z-scores from 5-300 to below 2 (below the detection threshold of 4) and reduces TPR at 5% FPR to below 50%, while improving quality by 27.5% and running 6 times faster than the best baseline on the long sequence generation. Our results suggest that robust AI-text detection via watermarking requires either accepting this fundamental vulnerability or unprecedented coordination among model providers.

📄 PDF Abstract BibTeX arXiv:2605.30501

Code (0)

등록된 구현이 없습니다.

Tasks

Text Detection

Similar Papers 제목 키워드 기반

More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles

2026-02-12 · Ruibo Chen, Yihan Wu, Xuehao Cui, Jingqi Zhang 외 arxiv

Watermarking has emerged as a crucial technique for detecting and attributing content generated by large language models. While recent advancements have utilized watermark ensembles to enhance robustness, prevailing meth…

FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks

2025-04-13 · Tianyi Wang, Harry Cheng, Ming-Hui Liu, Mohan Kankanhalli

Proactive Deepfake detection via robust watermarks has been raised ever since passive Deepfake detectors encountered challenges in identifying high-quality synthetic images. However, while demonstrating reasonable detect…

DeepFake DetectionFace Swapping

FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact

2026-08-04 · Alex Kwon arxiv

AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it …

Participation is not a Design Fix for Machine Learning

2020-07-05 · Mona Sloane, Emanuel Moss, Olaitan Awomolo, Laura Forlano

This paper critically examines existing modes of participation in design practice and machine learning. Cautioning against 'participation-washing', it suggests that the ML community must become attuned to possibly exploi…

BIG-bench Machine Learning

On the Effectiveness of Visible Watermarks

2017-07-01 · CVPR 2017 7 · Tali Dekel, Michael Rubinstein, Ce Liu, William T. Freeman

Visible watermarking is a widely-used technique for marking and protecting copyrights of many millions of images on the web, yet it suffers from an inherent security flaw---watermarks are typically added in a consistent …

Image Matting