paper-with-me

홈 › Papers

VICTOR: Dataset Copyright Auditing in Video Recognition Systems

2025-12-16 · Quan Yuan, Zhikun Zhang, Linkang Du, Min Chen, Mingyang Sun, Yunjun Gao, Shibo He, Jiming Chen arxiv

Video recognition systems are increasingly being deployed in daily life, such as content recommendation and security monitoring. To enhance video recognition development, many institutions have released high-quality public datasets with open-source licenses for training advanced models. At the same time, these datasets are also susceptible to misuse and infringement. Dataset copyright auditing is an effective solution to identify such unauthorized use. However, existing dataset copyright solutions primarily focus on the image domain; the complex nature of video data leaves dataset copyright auditing in the video domain unexplored. Specifically, video data introduces an additional temporal dimension, which poses significant challenges to the effectiveness and stealthiness of existing methods. In this paper, we propose VICTOR, the first dataset copyright auditing approach for video recognition systems. We develop a general and stealthy sample modification strategy that enhances the output discrepancy of the target model. By modifying only a small proportion of samples (e.g., 1%), VICTOR amplifies the impact of published modified samples on the prediction behavior of the target models. Then, the difference in the model's behavior for published modified and unpublished original samples can serve as a key basis for dataset auditing. Extensive experiments on multiple models and datasets highlight the superiority of VICTOR. Finally, we show that VICTOR is robust in the presence of several perturbation mechanisms to the training videos or the target models.

📄 PDF Abstract BibTeX arXiv:2512.14439

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SoK: Dataset Copyright Auditing in Machine Learning Systems

2024-10-22 · Linkang Du, Xuanru Zhou, Min Chen, Chusong Zhang 외

As the implementation of machine learning (ML) systems becomes more widespread, especially with the introduction of larger ML models, we perceive a spring demand for massive data. However, it inevitably causes infringeme…

Disguised Copyright Infringement of Latent Diffusion Models

2024-04-10 · Yiwei Lu, Matthew Y. R. Yang, Zuoqiu Liu, Gautam Kamath 외

Copyright infringement may occur when a generative model produces samples substantially similar to some copyrighted data that it had access to during the training phase. The notion of access usually refers to including c…

Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code

2024-02-14 · Vahid Majdinasab, Amin Nikanjam, Foutse khomh

Code auditing ensures that the developed code adheres to standards, regulations, and copyright protection by verifying that it does not contain code from protected sources. The recent advent of Large Language Models (LLM…

Clone Detection

Efficient Computation of Exact IRV Margins

2015-08-20 · Michelle Blom, Peter J. Stuckey, Vanessa J. Teague, Ron Tidhar

The margin of victory is easy to compute for many election schemes but difficult for Instant Runoff Voting (IRV). This is important because arguments about the correctness of an election outcome usually rely on the size …

SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking

2025-11-05 · Wenyuan Yang, Yichen Sun, Changzheng Chen, Zhixuan Chu 외 arxiv

Large-scale vision-language models, especially CLIP, have demonstrated remarkable performance across diverse downstream tasks. Soft prompts, as carefully crafted modules that efficiently adapt vision-language models to s…