paper-with-me

Papers

Efficiently Identifying Watermarked Segments in Mixed-Source Texts

2024-10-04 · Xuandong Zhao, Chenwen Liao, Yu-Xiang Wang, Lei LI

Text watermarks in large language models (LLMs) are increasingly used to detect synthetic text, mitigating misuse cases like fake news and academic dishonesty. While existing watermarking detection techniques primarily focus on classifying entire documents as watermarked or not, they often neglect the common scenario of identifying individual watermark segments within longer, mixed-source documents. Drawing inspiration from plagiarism detection systems, we propose two novel methods for partial watermark detection. First, we develop a geometry cover detection framework aimed at determining whether there is a watermark segment in long text. Second, we introduce an adaptive online learning algorithm to pinpoint the precise location of watermark segments within the text. Evaluated on three popular watermarking techniques (KGW-Watermark, Unigram-Watermark, and Gumbel-Watermark), our approach achieves high accuracy, significantly outperforming baseline methods. Moreover, our framework is adaptable to other watermarking techniques, offering new insights for precise watermark detection.

📄 PDF Abstract BibTeX arXiv:2410.03600

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents

2024-09-08 · Leyi Pan, Aiwei Liu, Yijian Lu, Zitian Gao 외

Watermarking algorithms for large language models (LLMs) have attained high accuracy in detecting LLM-generated text. However, existing methods primarily focus on distinguishing fully watermarked text from non-watermarke…

Computational EfficiencyText Detection

Baselines for Identifying Watermarked Large Language Models

2023-05-29 · Leonard Tang, Gavin Uberti, Tom Shlomi

We consider the emerging problem of identifying the presence and use of watermarking schemes in widely used, publicly hosted, closed source large language models (LLMs). We introduce a suite of baseline algorithms for id…

Watermark Anything with Localized Messages

2024-11-11 · Tom Sander, Pierre Fernandez, Alain Durmus, Teddy Furon 외

Image watermarking methods are not tailored to handle small watermarked areas. This restricts applications in real-world scenarios where parts of the image may come from different sources or have been edited. We introduc…

Fast segmentation of watermarked texts from large language models through an epidemic change-point framework

2025-09-25 · Soham Bonnerjee, Subhrajyoty Roy, Sayar Karmakar arxiv

With the growing use of large language models, concerns over content authenticity have spurred a variety of watermarking schemes. These schemes use secret keys to detect machine-generated text while remaining imperceptib…

Multi-Graph Decoding for Code-Switching ASR

2019-06-18 · Emre Yilmaz, Samuel Cohen, Xianghu Yue, David van Leeuwen 외

In the FAME! Project, a code-switching (CS) automatic speech recognition (ASR) system for Frisian-Dutch speech is developed that can accurately transcribe the local broadcaster's bilingual archives with CS speech. This a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modellingspeech-recognition+1