paper-with-me

홈 › Papers

Tracing Text Provenance via Context-Aware Lexical Substitution

2021-12-15 · Xi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang, Zehua Ma, Feng Wang, Nenghai Yu

Text content created by humans or language models is often stolen or misused by adversaries. Tracing text provenance can help claim the ownership of text content or identify the malicious users who distribute misleading content like machine-generated fake news. There have been some attempts to achieve this, mainly based on watermarking techniques. Specifically, traditional text watermarking methods embed watermarks by slightly altering text format like line spacing and font, which, however, are fragile to cross-media transmissions like OCR. Considering this, natural language watermarking methods represent watermarks by replacing words in original sentences with synonyms from handcrafted lexical resources (e.g., WordNet), but they do not consider the substitution's impact on the overall sentence's meaning. Recently, a transformer-based network was proposed to embed watermarks by modifying the unobtrusive words (e.g., function words), which also impair the sentence's logical and semantic coherence. Besides, one well-trained network fails on other different types of text content. To address the limitations mentioned above, we propose a natural language watermarking scheme based on context-aware lexical substitution (LS). Specifically, we employ BERT to suggest LS candidates by inferring the semantic relatedness between the candidates and the original sentence. Based on this, a selection strategy in terms of synchronicity and substitutability is further designed to test whether a word is exactly suitable for carrying the watermark signal. Extensive experiments demonstrate that, under both objective and subjective metrics, our watermarking scheme can well preserve the semantic integrity of original sentences and has a better transferability than existing methods. Besides, the proposed LS approach outperforms the state-of-the-art approach on the Stanford Word Substitution Benchmark.

📄 PDF Abstract BibTeX arXiv:2112.07873

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character Recognition (OCR)Sentence

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

RAPID: Robust APT Detection and Investigation Using Context-Aware Deep Learning

2024-06-08 · Yonatan Amaru, Prasanna Wudali, Yuval Elovici, Asaf Shabtai

Advanced persistent threats (APTs) pose significant challenges for organizations, leading to data breaches, financial losses, and reputational damage. Existing provenance-based approaches for APT detection often struggle…

Anomaly DetectionComputational Efficiency

From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents

2026-06-03 · Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu 외 arxiv

Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, retrieval, memory access, environmental interaction, and multi-agent collaboration.…

Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation

2025-12-24 · Kaiyuan Liu, Shaotian Yan, Rui Miao, Bing Wang 외 arxiv

Reasoning distillation has attracted increasing attention. It typically leverages a large teacher model to generate reasoning paths, which are then used to fine-tune a student model so that it mimics the teacher's behavi…

Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID

2025-08-27 · Xia Han, Qi Li, Jianbing Ni, Mohammad Zulkernine arxiv

Recent advances in LLM watermarking methods such as SynthID-Text by Google DeepMind offer promising solutions for tracing the provenance of AI-generated text. However, our robustness assessment reveals that SynthID-Text …

Information Retrieval

ScrollTimes: Tracing the Provenance of Paintings as a Window into History

2023-06-15 · Wei zhang, Wong Kam-Kwai, Yitian Chen, Ailing Jia 외

The study of cultural artifact provenance, tracing ownership and preservation, holds significant importance in archaeology and art history. Modern technology has advanced this field, yet challenges persist, including rec…