paper-with-me

홈 › Papers

Evading Data Provenance in Deep Neural Networks

2025-08-01 · Hongyu Zhu, Sichu Liang, Wenwen Wang, Zhuomeng Zhang, Fangqi Li, Shi-Lin Wang arxiv

Modern over-parameterized deep models are highly data-dependent, with large scale general-purpose and domain-specific datasets serving as the bedrock for rapid advancements. However, many datasets are proprietary or contain sensitive information, making unrestricted model training problematic. In the open world where data thefts cannot be fully prevented, Dataset Ownership Verification (DOV) has emerged as a promising method to protect copyright by detecting unauthorized model training and tracing illicit activities. Due to its diversity and superior stealth, evading DOV is considered extremely challenging. However, this paper identifies that previous studies have relied on oversimplistic evasion attacks for evaluation, leading to a false sense of security. We introduce a unified evasion framework, in which a teacher model first learns from the copyright dataset and then transfers task-relevant yet identifier-independent domain knowledge to a surrogate student using an out-of-distribution (OOD) dataset as the intermediary. Leveraging Vision-Language Models and Large Language Models, we curate the most informative and reliable subsets from the OOD gallery set as the final transfer set, and propose selectively transferring task-oriented knowledge to achieve a better trade-off between generalization and evasion effectiveness. Experiments across diverse datasets covering eleven DOV methods demonstrate our approach simultaneously eliminates all copyright identifiers and significantly outperforms nine state-of-the-art evasion attacks in both generalization and effectiveness, with moderate computational overhead. As a proof of concept, we reveal key vulnerabilities in current DOV methods, highlighting the need for long-term development to enhance practicality.

📄 PDF Abstract BibTeX arXiv:2508.01074

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors

2025-03-10 · Jingyi Zheng, Junfeng Wang, Zhen Sun, Wenhan Dong 외

As Large Language Models (LLMs) advance, Machine-Generated Texts (MGTs) have become increasingly fluent, high-quality, and informative. Existing wide-range MGT detectors are designed to identify MGTs to prevent the sprea…

Misinformation

EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System

2025-09-06 · Pavan Reddy, Aditya Sanjay Gujral arxiv

Large language model (LLM) assistants are increasingly integrated into enterprise workflows, raising new security concerns as they bridge internal and external data sources. This paper presents an in-depth case study of …

Provenance Graph Kernel

2020-10-20 · David Kohan Marzagão, Trung Dong Huynh, Ayah Helal, Sean Baccas 외

Provenance is a record that describes how entities, activities, and agents have influenced a piece of data; it is commonly represented as graphs with relevant labels on both their nodes and edges. With the growing adopti…

Provenance for the Description Logic ELHr

2020-01-21 · Camille Bourgaux, Ana Ozaki, Rafael Peñaloza, Livia Predoiu

We address the problem of handling provenance information in ELHr ontologies. We consider a setting recently introduced for ontology-based data access, based on semirings and extending classical data provenance, in which…

Provable Model Provenance Set for Large Language Models

2026-01-31 · Xiaoqi Qiu, Hao Zeng, Zhiyu Hou, Hongxin Wei arxiv

The growing prevalence of unauthorized model usage and misattribution has increased the need for reliable model provenance analysis. However, existing methods largely rely on heuristic fingerprint-matching rules that lac…