paper-with-me

홈 › Papers

Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data Hiding

2020-09-07 · Sahar Abdelnabi, Mario Fritz

Recent advances in natural language generation have introduced powerful language models with high-quality output text. However, this raises concerns about the potential misuse of such models for malicious purposes. In this paper, we study natural language watermarking as a defense to help better mark and trace the provenance of text. We introduce the Adversarial Watermarking Transformer (AWT) with a jointly trained encoder-decoder and adversarial training that, given an input text and a binary message, generates an output text that is unobtrusively encoded with the given message. We further study different training and inference strategies to achieve minimal changes to the semantics and correctness of the input text. AWT is the first end-to-end model to hide data in text by automatically learning -- without ground truth -- word substitutions along with their locations in order to encode the message. We empirically show that our model is effective in largely preserving text utility and decoding the watermark while hiding its presence against adversaries. Additionally, we demonstrate that our method is robust against a range of attacks.

📄 PDF Abstract BibTeX arXiv:2009.03015

Code (1)

S-Abdelnabi/awt pytorch

Tasks

DecoderDenoisingText Generation

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Tracing Text Provenance via Context-Aware Lexical Substitution

2021-12-15 · Xi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang 외

Text content created by humans or language models is often stolen or misused by adversaries. Tracing text provenance can help claim the ownership of text content or identify the malicious users who distribute misleading …

Optical Character Recognition (OCR)Sentence

Data Provenance for Image Auto-Regressive Generation

2026-06-22 · Bihe Zhao, Louis Kerner, Michel Meintz, Tameem Bakr 외 arxiv

Image autoregressive models (IARs) have recently demonstrated remarkable capabilities in visual content generation, achieving photorealistic quality and rapid synthesis through the next-token prediction paradigm adapted …

Image Generation

Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID

2025-08-27 · Xia Han, Qi Li, Jianbing Ni, Mohammad Zulkernine arxiv

Recent advances in LLM watermarking methods such as SynthID-Text by Google DeepMind offer promising solutions for tracing the provenance of AI-generated text. However, our robustness assessment reveals that SynthID-Text …

Information Retrieval

Pixel Seal: Adversarial-only training for invisible image and video watermarking

2025-12-18 · Tomáš Souček, Pierre Fernandez, Hady Elsahar, Sylvestre-Alvise Rebuffi 외 arxiv

Invisible watermarking is essential for tracing the provenance of digital content. However, training state-of-the-art models remains notoriously difficult, with current approaches often struggling to balance robustness a…

MC$^2$Mark: Distortion-Free Multi-Bit Watermarking for Long Messages

2026-02-15 · Xuehao Cui, Ruibo Chen, Yihan Wu, Heng Huang arxiv

Large language models now produce text indistinguishable from human writing, which increases the need for reliable provenance tracing. Multi-bit watermarking can embed identifiers into generated text, but existing method…