paper-with-me

홈 › Papers

TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection

2026-05-12 · Tom Sander, Hongyan Chang, Tomáš Souček, Tuan Tran, Valeriu Lacatusu, Sylvestre-Alvise Rebuffi, Alexandre Mourachko, Surya Parimi, Christophe Ropers, Rashel Moritz, Vanessa Stark, Hady Elsahar, Pierre Fernandez arxiv

We introduce TextSeal, a state-of-the-art watermark for large language models. Building on Gumbel-max sampling, TextSeal introduces dual-key generation to restore output diversity, along with entropy-weighted scoring and multi-region localization for improved detection. It supports serving optimizations such as speculative decoding and multi-token prediction, and does not add any inference overhead. TextSeal strictly dominates baselines like SynthID-text in detection strength and is robust to dilution, maintaining confident localized detection even in heavily mixed human/AI documents. The scheme is theoretically distortion-free, and evaluation across reasoning benchmarks confirms that it preserves downstream performance; while a multilingual human evaluation (6000 A/B comparisons, 5 languages) shows no perceptible quality difference. Beyond its use for provenance detection, TextSeal is also ``radioactive'': its watermark signal transfers through model distillation, enabling detection of unauthorized use.

📄 PDF Abstract BibTeX arXiv:2605.12456

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Geometry-Aware Localized Watermarking for Copyright Protection in Embedding-as-a-Service

2026-04-13 · Zhimin Chen, Xiaojie Liang, Wenbo Xu, Yuxuan Liu 외 arxiv

Embedding-as-a-Service (EaaS) has become an important semantic infrastructure for natural language and multimedia applications, but it is highly vulnerable to model stealing and copyright infringement. Existing EaaS wate…

High-Fidelity Face Content Recovery via Tamper-Resilient Versatile Watermarking

2026-03-25 · Peipeng Yu, Jinfeng Xie, Chengfu Ou, Xiaoyu Zhou 외 arxiv

The proliferation of AIGC-driven face manipulation and deepfakes poses severe threats to media provenance, integrity, and copyright protection. Existing versatile watermarking systems typically rely on embedding explicit…

Multi-Agent Framework for Controllable and Protected Generative Content Creation: Addressing Copyright and Provenance in AI-Generated Media

2026-01-09 · Haris Khan, Sadia Asif, Shumaila Asif arxiv

The proliferation of generative AI systems creates unprecedented opportunities for content creation while raising critical concerns about controllability, copyright infringement, and content provenance. Current generativ…

Keyed Provenance Watermarking with Complementary Lattice-Based Secure Aggregation for Federated Learning

2026-08-20 · Xinyun Liu, Zhi Lu, Yu Chen, Ronghua Xu arxiv

Federated learning (FL) is vulnerable to multi-level attacks. However, existing methods address them separately, leaving FL exposed to data leakage, unauthorized reuse, and malicious gradient manipulation. In this work, …

Federated Learning

On the Coexistence and Ensembling of Watermarks

2025-01-29 · Aleksandar Petrov, Shruti Agarwal, Philip H. S. Torr, Adel Bibi 외

Watermarking, the practice of embedding imperceptible information into media such as images, videos, audio, and text, is essential for intellectual property protection, content provenance and attribution. The growing com…