paper-with-me

홈 › Papers

Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models

2025-12-15 · Malte Hellmeier arxiv

Securing digital text is becoming increasingly relevant due to the widespread use of large language models. Individuals' fear of losing control over data when it is being used to train such machine learning models or when distinguishing model-generated output from text written by humans. Digital watermarking provides additional protection by embedding an invisible watermark within the data that requires protection. However, little work has been taken to analyze and verify if existing digital text watermarking methods are secure and undetectable by large language models. In this paper, we investigate the security-related area of watermarking and machine learning models for text data. In a controlled testbed of three experiments, ten existing Unicode text watermarking methods were implemented and analyzed across six large language models: GPT-5, GPT-4o, Teuken 7B, Llama 3.3, Claude Sonnet 4, and Gemini 2.5 Pro. The findings of our experiments indicate that, especially the latest reasoning models, can detect a watermarked text. Nevertheless, all models fail to extract the watermark unless implementation details in the form of source code are provided. We discuss the implications for security researchers and practitioners and outline future research opportunities to address security concerns.

📄 PDF Abstract BibTeX arXiv:2512.13325

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WaterSearch: Exploring Seed Pooling for Improving the Quality-Detectability Trade-off in LLM Watermarking

2025-11-30 · Yukang Lin, Jiahao Shao, Shuoran Jiang, Wentao Zhu 외 arxiv

Watermarking acts as a critical safeguard in text generated by Large Language Models (LLMs). By embedding identifiable signals into model outputs, watermarking enables reliable attribution and enhances the security of ma…

Text Generation

Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning

2025-04-09 · Li An, Yujian Liu, Yepeng Liu, Yang Zhang 외

Watermarking has emerged as a promising technique for detecting texts generated by LLMs. Current research has primarily focused on three design criteria: high quality of the watermarked text, high detectability, and robu…

Representation Learning

Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization

2025-10-13 · Chenrui Wang, Junyi Shu, Billy Chiu, Yu Li 외 arxiv

The rapid development of LLMs has raised concerns about their potential misuse, leading to various watermarking schemes that typically offer high detectability. However, existing watermarking techniques often face trade-…

From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models

2025-05-15 · Yidan Wang, Yubing Ren, Yanan Cao, Binxing Fang

The rise of Large Language Models (LLMs) has heightened concerns about the misuse of AI-generated text, making watermarking a promising solution. Mainstream watermarking schemes for LLMs fall into two categories: logits-…

Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models

2024-02-28 · Mingjia Huo, Sai Ashish Somayajula, Youwei Liang, Ruisi Zhang 외

Large language models generate high-quality responses with potential misinformation, underscoring the need for regulation by distinguishing AI-generated and human-written texts. Watermarking is pivotal in this context, w…

Misinformation