paper-with-me

홈 › Papers

How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets

2025-11-10 · Chung Peng Lee, Rachel Hong, Harry H. Jiang, Aster Plotnik, William Agnew, Jamie Morgenstern arxiv

The internet has become the main source of data to train modern text-to-image or vision-language models, yet it is increasingly unclear whether web-scale data collection practices for training AI systems adequately respect data owners' wishes. Ignoring the owner's indication of consent around data usage not only raises ethical concerns but also has recently been elevated into lawsuits around copyright infringement cases. In this work, we aim to reveal information about data owners' consent to AI scraping and training, and study how it's expressed in DataComp, a popular dataset of 12.8 billion text-image pairs. We examine both the sample-level information, including the copyright notice, watermarking, and metadata, and the web-domain-level information, such as a site's Terms of Service (ToS) and Robots Exclusion Protocol. We estimate at least 122M of samples exhibit some indication of copyright notice in CommonPool, and find that 60\% of the samples in the top 50 domains come from websites with ToS that prohibit scraping. Furthermore, we estimate 9-13\% with 95\% confidence interval of samples from CommonPool to contain watermarks, where existing watermark detection methods fail to capture them in high fidelity. Our holistic methods and findings show that data owners rely on various channels to convey data consent, of which current AI data collection pipelines do not entirely respect. These findings highlight the limitations of the current dataset curation/release practice and the need for a unified data consent framework taking AI purposes into consideration.

📄 PDF Abstract BibTeX arXiv:2511.08637

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Yes, But Not Always. Generative AI Needs Nuanced Opt-in

2026-04-10 · Wiebke Hutiri, Morgan Scheuerman, Shruti Nagpal, Austin Hoag 외 arxiv

This paper argues that a one-size-fits-all approach to specifying consent for the use of creative works in generative AI is insufficient. Real-world ownership and rights holder structures, the imitation of artistic style…

NLP--Based Readability Assessment of Health--Related Texts: a Case Study on Italian Informed Consent Forms

2015-09-01 · WS 2015 9 · Giulia Venturi, Bell, Tommaso i, Felice Dell{'}Orletta 외

Welfare v. Consent: On the Optimal Penalty for Harassment

2021-03-01 · Ratul Das Chaudhury, Birendra Rai, Liang Choon Wang, Dyuti Banerjee

The economic approach to determine optimal legal policies involves maximizing a social welfare function. We propose an alternative: a consent-approach that seeks to promote consensual interactions and deter non-consensua…

DECORAIT -- DECentralized Opt-in/out Registry for AI Training

2023-09-25 · Kar Balan, Alex Black, Simon Jenni, Andrew Gilbert 외

We present DECORAIT; a decentralized registry through which content creators may assert their right to opt in or out of AI training as well as receive reward for their contributions. Generative AI (GenAI) enables images …

Inform the uninformed: Improving Online Informed Consent Reading with an AI-Powered Chatbot

2023-02-02 · Ziang Xiao, Tiffany Wenting Li, Karrie Karahalios, Hari Sundaram

Informed consent is a core cornerstone of ethics in human subject research. Through the informed consent process, participants learn about the study procedure, benefits, risks, and more to make an informed decision. Howe…

ChatbotEthicsForm