paper-with-me

홈 › Papers

Neglected Risks: The Disturbing Reality of Children's Images in Datasets and the Urgent Call for Accountability

2025-04-20 · Carlos Caetano, Gabriel O. dos Santos, Caio Petrucci, Artur Barros, Camila Laranjeira, Leo S. F. Ribeiro, Júlia F. de Mendonça, Jefersson A. dos Santos, Sandra Avila

Including children's images in datasets has raised ethical concerns, particularly regarding privacy, consent, data protection, and accountability. These datasets, often built by scraping publicly available images from the Internet, can expose children to risks such as exploitation, profiling, and tracking. Despite the growing recognition of these issues, approaches for addressing them remain limited. We explore the ethical implications of using children's images in AI datasets and propose a pipeline to detect and remove such images. As a use case, we built the pipeline on a Vision-Language Model under the Visual Question Answering task and tested it on the #PraCegoVer dataset. We also evaluate the pipeline on a subset of 100,000 images from the Open Images V7 dataset to assess its effectiveness in detecting and removing images of children. The pipeline serves as a baseline for future research, providing a starting point for more comprehensive tools and methodologies. While we leverage existing models trained on potentially problematic data, our goal is to expose and address this issue. We do not advocate for training or deploying such models, but instead call for urgent community reflection and action to protect children's rights. Ultimately, we aim to encourage the research community to exercise - more than an additional - care in creating new datasets and to inspire the development of tools to protect the fundamental rights of vulnerable groups, particularly children.

📄 PDF Abstract BibTeX arXiv:2504.14446

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

Combating the Elsagate phenomenon: Deep learning architectures for disturbing cartoons

2019-04-18 · Akari Ishikawa, Edson Bollis, Sandra Avila

Watching cartoons can be useful for children's intellectual, social and emotional development. However, the most popular video sharing platform today provides many videos with Elsagate content. Elsagate is a phenomenon t…

Deep LearningPornography Detection

Malicious or Benign? Towards Effective Content Moderation for Children's Videos

2023-05-24 · Syed Hammad Ahmed, Muhammad Junaid Khan, H. M. Umer Qaisar, Gita Sukthankar

Online video platforms receive hundreds of hours of uploads every minute, making manual content moderation impossible. Unfortunately, the most vulnerable consumers of malicious video content are children from ages 1-5 wh…

Video Classification

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

2026-06-24 · Minh-Kha Nguyen, Trung-Hieu Do, Kim Anh Phung, Thao Thi Phuong Dao 외 arxiv

Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lacking parental supervision. Identifying actions that pose risks plays a …

Action Recognition

Age-Conditioned Synthesis of Pediatric Computed Tomography with Auxiliary Classifier Generative Adversarial Networks

2020-01-31 · Chi Nok Enoch Kan, Najibakram Maheenaboobacker, Dong Hye Ye

Deep learning is a popular and powerful tool in computed tomography (CT) image processing such as organ segmentation, but its requirement of large training datasets remains a challenge. Even though there is a large anato…

Computed Tomography (CT)Generative Adversarial NetworkOrgan Segmentation

MinorBench: A hand-built benchmark for content-based risks for children

2025-03-13 · Shaun Khoo, Gabriel Chua, Rachel Shong

Large Language Models (LLMs) are rapidly entering children's lives - through parent-driven adoption, schools, and peer networks - yet current AI ethics and safety research do not adequately address content-related risks …

ChatbotEthics