paper-with-me

홈 › Papers

HoneyImage: Verifiable, Harmless, and Stealthy Dataset Ownership Verification for Image Models

2025-07-27 · Zhihao Zhu, Jiale Han, Yi Yang arxiv

Image-based AI models are increasingly deployed across a wide range of domains, including healthcare, security, and consumer applications. However, many image datasets carry sensitive or proprietary content, raising critical concerns about unauthorized data usage. Data owners therefore need reliable mechanisms to verify whether their proprietary data has been misused to train third-party models. Existing solutions, such as backdoor watermarking and membership inference, face inherent trade-offs between verification effectiveness and preservation of data integrity. In this work, we propose HoneyImage, a novel method for dataset ownership verification in image recognition models. HoneyImage selectively modifies a small number of hard samples to embed imperceptible yet verifiable traces, enabling reliable ownership verification while maintaining dataset integrity. Extensive experiments across four benchmark datasets and multiple model architectures show that HoneyImage consistently achieves strong verification accuracy with minimal impact on downstream performance while maintaining imperceptible. The proposed HoneyImage method could provide data owners with a practical mechanism to protect ownership over valuable image datasets, encouraging safe sharing and unlocking the full transformative potential of data-driven AI.

📄 PDF Abstract BibTeX arXiv:2508.00892

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided Backdoors

2026-01-23 · Qinkai Yu, Chong Zhang, Gaojie Jin, Tianjin Huang 외 arxiv

Annotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associa…

Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models

2026-05-09 · Ming Sun, Rui Wang, Xingrui Yu, Lihua Jing 외 arxiv

Vision-Language-Action models (VLAs) support generalist robotic control by enabling end-to-end decision policies directly from multi-modal inputs. As trained VLAs are increasingly shared and adapted, protecting model own…

Data Taggants: Dataset Ownership Verification via Harmless Targeted Data Poisoning

2024-10-09 · Wassim Bouaziz, El-Mahdi El-Mhamdi, Nicolas Usunier

Dataset ownership verification, the process of determining if a dataset is used in a model's training data, is necessary for detecting unauthorized data usage and data contamination. Existing approaches, such as backdoor…

Data Poisoning

Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright Protection

2022-09-27 · Yiming Li, Yang Bai, Yong Jiang, Yong Yang 외

Deep neural networks (DNNs) have demonstrated their superiority in practice. Arguably, the rapid development of DNNs is largely benefited from high-quality (open-sourced) datasets, based on which researchers and develope…

EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright Protection

2025-01-01 · CVPR 2025 1 · Ming Sun, Rui Wang, Zixuan Zhu, Lihua Jing 외

High-quality open-source datasets are essential for advancing deep neural networks. However, the unauthorized commercial use of these datasets has raised significant concerns about copyright protection. One promising…

Prediction