paper-with-me

홈 › Papers

Learnable Linguistic Watermarks for Tracing Model Extraction Attacks on Large Language Models

2024-04-28 · Minhao Bai, Kaiyi Pang, Yongfeng Huang

In the rapidly evolving domain of artificial intelligence, safeguarding the intellectual property of Large Language Models (LLMs) is increasingly crucial. Current watermarking techniques against model extraction attacks, which rely on signal insertion in model logits or post-processing of generated text, remain largely heuristic. We propose a novel method for embedding learnable linguistic watermarks in LLMs, aimed at tracing and preventing model extraction attacks. Our approach subtly modifies the LLM's output distribution by introducing controlled noise into token frequency distributions, embedding an statistically identifiable controllable watermark.We leverage statistical hypothesis testing and information theory, particularly focusing on Kullback-Leibler Divergence, to differentiate between original and modified distributions effectively. Our watermarking method strikes a delicate well balance between robustness and output quality, maintaining low false positive/negative rates and preserving the LLM's original performance.

📄 PDF Abstract BibTeX arXiv:2405.01509

Code (0)

등록된 구현이 없습니다.

Tasks

Model extraction

Similar Papers 제목 키워드 기반

On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks

2024-07-05 · Zesen Liu, Tianshuo Cong, Xinlei He, Qi Li

Large Language Models (LLMs) excel in various applications, including text generation and complex tasks. However, the misuse of LLMs raises concerns about the authenticity and ethical implications of the content they pro…

Face SwappingText Generation

DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks

2025-11-12 · Yunfei Yang, Xiaojun Chen, Yuexin Xuan, Zhendong Zhao 외 arxiv

Model watermarking techniques can embed watermark information into the protected model for ownership declaration by constructing specific input-output pairs. However, existing watermarks are easily removed when facing mo…

T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking

2026-06-10 · Jian-Ping Mei, Weibin Zhang, Ao Yao, Tiantian Zhu 외 arxiv

Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against var…

Model extraction

Was my Model Stolen? Feature Sharing for Robust and Transferable Watermarks

2021-09-29 · Ruixiang Tang, Hongye Jin, Curtis Wigington, Mengnan Du 외

Deep Neural Networks (DNNs) are increasingly being deployed in cloud-based services via various APIs, e.g., prediction APIs. Recent studies show that these public APIs are vulnerable to the model extraction attack, where…

Model extraction

WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection

2024-03-03 · Anudeex Shetty, Yue Teng, Ke He, Qiongkai Xu

Embedding as a Service (EaaS) has become a widely adopted solution, which offers feature extraction capabilities for addressing various downstream tasks in Natural Language Processing (NLP). Prior studies have shown that…

Model extraction