paper-with-me

홈 › Papers

MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking

2025-12-03 · Yizhou Zhao, Zhiwei Steven Wu, Adam Block arxiv

Watermarking aims to embed hidden signals in generated text that can be reliably detected when given access to a secret key. Open-weight language models pose acute challenges for such watermarking schemes because the inference-time interventions that dominate contemporary approaches cannot be enforced once model weights are public. Existing watermaking techniques for open-weight models, such as the recently proposed GaussMark, typically rely on small modifications to model weights, which can yield signals detectable to those equipped with a secret key, but achieving detection power comparable to inference-time watermarks generally requires weight perturbations that noticeably reduce generation quality. We introduce MarkTune, a theoretically principled, on-policy fine-tuning framework that treats the GaussMark signal as a reward while simultaneously regularizing against degradation in text quality. We derive MarkTune as an improvement on GaussMark and demonstrate that MarkTune consistently improves the quality-detectability trade-off over GaussMark by steering finer-grained, watermark-aware weight updates within the model's representation space while preserving generation quality. Empirically, we show that MarkTune pushes the quality-detectability frontier of GaussMark close to that of inference-time watermarking, remains robust to paraphrasing and fine-tuning attacks, and exhibits strong generalization: a model fine-tuned on one dataset retains substantial watermark detection power on unseen datasets. Together, these results establish MarkTune as a general strategy for embedding robust, high-quality watermarks into open-weight LMs.

📄 PDF Abstract BibTeX arXiv:2512.04044

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward Stronger Code Watermarking: A Grammar-Driven Approach to Optimizing the Trade-off Between Quality and Detectability

2026-07-11 · Licheng Yu, Aiwei Liu, Songze Li arxiv

With the rapid development of Large Language Models (LLMs), text watermarking has emerged as a crucial technique for identifying machine-generated content. However, directly applying existing logits-based watermarking me…

Code Generation

WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off

2024-03-06 · Eva Giboulot, Teddy Furon

Watermarking is a technical means to dissuade malfeasant usage of Large Language Models. This paper proposes a novel watermarking scheme, so-called WaterMax, that enjoys high detectability while sustaining the quality of…

Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization

2025-10-13 · Chenrui Wang, Junyi Shu, Billy Chiu, Yu Li 외 arxiv

The rapid development of LLMs has raised concerns about their potential misuse, leading to various watermarking schemes that typically offer high detectability. However, existing watermarking techniques often face trade-…

PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs

2025-10-27 · Jiaqi Xue, Yifei Zhao, Mansour Al Ghanim, Shangqian Gao 외 arxiv

Text watermarking for large language models (LLMs) enables model owners to verify text origin and protect intellectual property. While watermarking methods for closed-source LLMs are relatively mature, extending them to …

Medical Image Quality Metrics for Foveated Model Observers

2021-02-09 · Miguel A. Lago, Craig K. Abbey, Miguel P. Eckstein

A recently proposed model observer mimics the foveated nature of the human visual system by processing the entire image with varying spatial detail, executing eye movements and scrolling through slices. The model can pre…

model