paper-with-me

Papers

Baselines for Identifying Watermarked Large Language Models

2023-05-29 · Leonard Tang, Gavin Uberti, Tom Shlomi

We consider the emerging problem of identifying the presence and use of watermarking schemes in widely used, publicly hosted, closed source large language models (LLMs). We introduce a suite of baseline algorithms for identifying watermarks in LLMs that rely on analyzing distributions of output tokens and logits generated by watermarked and unmarked LLMs. Notably, watermarked LLMs tend to produce distributions that diverge qualitatively and identifiably from standard models. Furthermore, we investigate the identifiability of watermarks at varying strengths and consider the tradeoffs of each of our identification mechanisms with respect to watermarking scenario. Along the way, we formalize the specific problem of identifying watermarks in LLMs, as well as LLM watermarks and watermark detection in general, providing a framework and foundations for studying them.

📄 PDF Abstract BibTeX arXiv:2305.18456

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficiently Identifying Watermarked Segments in Mixed-Source Texts

2024-10-04 · Xuandong Zhao, Chenwen Liao, Yu-Xiang Wang, Lei LI

Text watermarks in large language models (LLMs) are increasingly used to detect synthetic text, mitigating misuse cases like fake news and academic dishonesty. While existing watermarking detection techniques primarily f…

Protecting Intellectual Property of Language Generation APIs with Lexical Watermark

2021-12-05 · Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu 외

Nowadays, due to the breakthrough in natural language generation (NLG), including machine translation, document summarization, image captioning, etc NLG models have been encapsulated in cloud APIs to serve over half a bi…

Document SummarizationImage CaptioningMachine TranslationModel extraction+1

I Know You Did Not Write That! A Sampling Based Watermarking Method for Identifying Machine Generated Text

2023-11-29 · Kaan Efe Keleş, Ömer Kaan Gürbüz, Mucahid Kutlu

Potential harms of Large Language Models such as mass misinformation and plagiarism can be partially mitigated if there exists a reliable way to detect machine generated text. In this paper, we propose a new watermarking…

Misinformation

WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents

2024-09-08 · Leyi Pan, Aiwei Liu, Yijian Lu, Zitian Gao 외

Watermarking algorithms for large language models (LLMs) have attained high accuracy in detecting LLM-generated text. However, existing methods primarily focus on distinguishing fully watermarked text from non-watermarke…

Computational EfficiencyText Detection

Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks

2023-09-29 · Mehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar 외

In light of recent advancements in generative AI models, it has become essential to distinguish genuine content from AI-generated one to prevent the malicious usage of fake materials as authentic ones and vice versa. Var…

Adversarial AttackFace Swapping