paper-with-me

홈 › Papers

On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical Validation

2026-03-03 · Romina Omidi, Yun Dong, Binghui Wang arxiv

Google's SynthID-Text, the first ever production-ready generative watermark system for large language model, designs a novel Tournament-based method that achieves the state-of-the-art detectability for identifying AI-generated texts. The system's innovation lies in: 1) a new Tournament sampling algorithm for watermarking embedding, 2) a detection strategy based on the introduced score function (e.g., Bayesian or mean score), and 3) a unified design that supports both distortionary and non-distortionary watermarking methods. This paper presents the first theoretical analysis of SynthID-Text, with a focus on its detection performance and watermark robustness, complemented by empirical validation. For example, we prove that the mean score is inherently vulnerable to increased tournament layers, and design a layer inflation attack to break SynthID-Text. We also prove the Bayesian score offers improved watermark robustness w.r.t. layers and further establish that the optimal Bernoulli distribution for watermark detection is achieved when the parameter is set to 0.5. Together, these theoretical and empirical insights not only deepen our understanding of SynthID-Text, but also open new avenues for analyzing effective watermark removal strategies and designing robust watermarking techniques. Source code is available at https: //github.com/romidi80/Synth-ID-Empirical-Analysis.

📄 PDF Abstract BibTeX arXiv:2603.03410

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID

2025-08-27 · Xia Han, Qi Li, Jianbing Ni, Mohammad Zulkernine arxiv

Recent advances in LLM watermarking methods such as SynthID-Text by Google DeepMind offer promising solutions for tracing the provenance of AI-generated text. However, our robustness assessment reveals that SynthID-Text …

Information Retrieval

SynthID-Image: Image watermarking at internet scale

2025-10-10 · Sven Gowal, Rudy Bunel, Florian Stimberg, David Stutz 외 arxiv

We introduce SynthID-Image, a deep learning-based system for invisibly watermarking AI-generated imagery. This paper documents the technical desiderata, threat models, and practical challenges of deploying such a system …

Watermarks Without Verification: AI Text Watermarking After the EU AI Act

2026-09-09 · Alexander Nemecek, Vipin Chaudhary, Erman Ayday arxiv

On August 2, 2026, the obligations of Article 50 of the EU AI Act took effect, requiring generative AI providers to mark the content their systems produce and ensure it can be detected as AI-generated. Days later, Anthro…

Topic-Based Watermarks for Large Language Models

2024-04-02 · Alexander Nemecek, Yuzhou Jiang, Erman Ayday

The indistinguishability of Large Language Model (LLM) output from human-authored content poses significant challenges, raising concerns about potential misuse of AI-generated text and its influence on future AI model tr…

Language ModelingLanguage ModellingLarge Language ModelText Generation

AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation

2026-07-17 · Saifur Rahman Tamim, Amir Labib Khan arxiv

Governments are increasingly mandating that LLM-generated content carry watermarks. The EU AI Act calls for markings that are "sufficiently reliable and robust." California's SB 942 requires disclosure that is "permanent…