Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking
Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-off between detectability and semantic distortion. Existing analyses provide limited guidance for principled hyperparameter selection, leaving practical deployments reliant on heuristic tuning. In this work, we develop a power-calibrated statistical framework that establishes explicit quantitative relationships between watermark hyperparameters, detection power, and distortion. This characterization transforms watermark design into a guided optimization problem. Building on these results, we derive practical parameter selection procedures that achieve optimal tradeoffs under constraints. Extensive experiments across multiple language models and datasets validate the theory and demonstrate that the proposed framework consistently identifies Pareto-optimal points.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Fine-tuning Is Not Enough: A Simple yet Effective Watermark Removal Attack for DNN Models
Watermarking has become the tendency in protecting the intellectual property of DNN models. Recent works, from the adversary's perspective, attempted to subvert watermarking mechanisms by designing watermark removal atta…
MemorizationA Locally Tokenized Generative Model for Robust Time-Series Watermarking
Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by reliability failures under post-editing attacks. We show that existing detectors, wh…
A Recipe for Watermarking Diffusion Models
Diffusion models (DMs) have demonstrated advantageous potential on generative tasks. Widespread interest exists in incorporating DMs into downstream applications, such as producing or editing photorealistic images. Howev…
MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
Watermarking aims to embed hidden signals in generated text that can be reliably detected when given access to a secret key. Open-weight language models pose acute challenges for such watermarking schemes because the inf…
Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
We show the viability of tackling misuses of large language models beyond the identification of machine-generated text. While existing zero-bit watermark methods focus on detection only, some malicious misuses demand tra…
Language ModelingLanguage ModellingPosition