paper-with-me

Papers

Data Leakage in Visual Datasets

2025-08-24 · Patrick Ramos, Ryan Ramos, Noa Garcia arxiv

We analyze data leakage in visual datasets. Data leakage refers to images in evaluation benchmarks that have been seen during training, compromising fair model evaluation. Given that large-scale datasets are often sourced from the internet, where many computer vision benchmarks are publicly available, our efforts are focused into identifying and studying this phenomenon. We characterize visual leakage into different types according to its modality, coverage, and degree. By applying image retrieval techniques, we unequivocally show that all the analyzed datasets present some form of leakage, and that all types of leakage, from severe instances to more subtle cases, compromise the reliability of model evaluation in downstream tasks.

📄 PDF Abstract BibTeX arXiv:2508.17416

Code (0)

등록된 구현이 없습니다.

Tasks

Image Retrieval

Similar Papers 제목 키워드 기반

Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets

2025-11-17 · Noam Glazner, Noam Tsfaty, Sharon Shalev, Avishai Weizman arxiv

We propose a cluster-based frame selection strategy to mitigate information leakage in video-derived frames datasets. By grouping visually similar frames before splitting into training, validation, and test sets, the met…

Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs

2026-08-13 · Beining Xu, Hairui Wang, Jiaxin Wang, Changsheng Chen 외 arxiv

While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding M…

Key Information Extraction

StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance

2025-10-08 · Jaeseok Jeong, Junho Kim, Gayoung Lee, Yunjey Choi 외 arxiv

In the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled more precise control over style and conten…

Text-to-Image Generation

From Measurement to Mitigation: Quantifying and Reducing Identity Leakage in Image Representation Encoders with Linear Subspace Removal

2026-04-07 · Daniel George, Charles Yeh, Daniel Lee, Yifei Zhang arxiv

Frozen visual embeddings (e.g., CLIP, DINOv2/v3, SSCD) power retrieval and integrity systems, yet their use on face-containing data is constrained by unmeasured identity leakage and a lack of deployable mitigations. We t…

Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios

2025-09-28 · Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu Li arxiv

Visual modality is the most vulnerable to privacy leakage in real-world multimodal applications like autonomous driving with visual and radar data; Machine unlearning removes specific training data from pre-trained model…

Contrastive LearningAutonomous Driving