paper-with-me

홈 › Papers

AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation

2026-07-17 · Saifur Rahman Tamim, Amir Labib Khan arxiv

Governments are increasingly mandating that LLM-generated content carry watermarks. The EU AI Act calls for markings that are "sufficiently reliable and robust." California's SB 942 requires disclosure that is "permanent or extraordinarily difficult to remove." Both mandates rest on an untested assumption: that watermark detection yields evidence reliable enough for courts. This paper tests that assumption directly. We evaluate three representative LLM watermarking methods -- KGW, Unigram, and the MarkLLM implementation of SynthID-Text -- against the Daubert admissibility criteria and the NIST SP 800-86 digital forensic process. To structure this evaluation, we propose a Forensic Readiness Score (FRS) framework with 12 criteria, three mandatory gates, and a 60-point scoring system. We focus on meaning-preserving paraphrase as the attack vector, since it is both legally realistic and difficult to dismiss as evidence tampering. The results raise serious evidentiary concerns. Out of 846 valid paraphrase runs across 15 diverse prompts per method, every single initially-detected KGW and Unigram text lost its watermark after paraphrasing -- 100% conditional removal. SynthID fared only slightly better at 98.3%. Even before any attack, false-negative rates were already high: 70% for KGW, 83% for Unigram, 80% for SynthID. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated and showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband. None of the three methods satisfy more than two of five Daubert factors. We also find that the FRS point-based scoring system, despite working as designed, cannot fully capture forensic uselessness -- a limitation worth noting for future framework design. These configurations, as tested, do not meet the evidentiary bar that courts require.

📄 PDF Abstract BibTeX arXiv:2607.16010

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are Watermarks Bugs for Deepfake Detectors? Rethinking Proactive Forensics

2024-04-27 · Xiaoshuai Wu, Xin Liao, Bo Ou, Yuling Liu 외

AI-generated content has accelerated the topic of media synthesis, particularly Deepfake, which can manipulate our portraits for positive or malicious purposes. Before releasing these threatening face images, one promisi…

DeepFake DetectionFace Swapping

The Forensic Cost of Watermark Removal: From Dedicated Attacks to Image Editing

2026-04-28 · Gautier Evennou, Ewa Kijak arxiv

Current watermark removal methods are evaluated on two axes: attack success rate and perceptual quality. We show this is insufficient. While state-of-the-art attacks successfully degrade the watermark signal without visi…

Image Editing

StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models

2025-09-22 · Haoxin Yang, Bangzhen Liu, Xuemiao Xu, Cheng Xu 외 arxiv

The advancement of diffusion models has enhanced the realism of AI-generated content but also raised concerns about misuse, necessitating robust copyright protection and tampering localization. Although recent methods ha…

High-Fidelity Face Content Recovery via Tamper-Resilient Versatile Watermarking

2026-03-25 · Peipeng Yu, Jinfeng Xie, Chengfu Ou, Xiaoyu Zhou 외 arxiv

The proliferation of AIGC-driven face manipulation and deepfakes poses severe threats to media provenance, integrity, and copyright protection. Existing versatile watermarking systems typically rely on embedding explicit…

Uncovering and Mitigating Destructive Multi-Embedding Attacks in Deepfake Proactive Forensics

2025-08-24 · Lixin Jia, Haiyang Sun, Zhiqing Guo, Yunfeng Diao 외 arxiv

With the rapid evolution of deepfake technologies and the wide dissemination of digital media, personal privacy is facing increasingly serious security threats. Deepfake proactive forensics, which involves embedding impe…