paper-with-me

홈 › Papers

Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP

2022-08-10 · Thao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh, Ludwig Schmidt

Web-crawled datasets have enabled remarkable generalization capabilities in recent image-text models such as CLIP (Contrastive Language-Image pre-training) or Flamingo, but little is known about the dataset creation processes. In this work, we introduce a testbed of six publicly available data sources - YFCC, LAION, Conceptual Captions, WIT, RedCaps, Shutterstock - to investigate how pre-training distributions induce robustness in CLIP. We find that the performance of the pre-training data varies substantially across distribution shifts, with no single data source dominating. Moreover, we systematically study the interactions between these data sources and find that combining multiple sources does not necessarily yield better models, but rather dilutes the robustness of the best individual data source. We complement our empirical findings with theoretical insights from a simple setting, where combining the training data also results in diluted robustness. In addition, our theoretical model provides a candidate explanation for the success of the CLIP-based data filtering technique recently employed in the LAION dataset. Overall our results demonstrate that simply gathering a large amount of data from the web is not the most effective way to build a pre-training dataset for robust generalization, necessitating further study into dataset design. Code is available at https://github.com/mlfoundations/clip_quality_not_quantity.

📄 PDF Abstract BibTeX arXiv:2208.05516

Code (1)

mlfoundations/clip_quality_not_quantity 공식 구현

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

A Case Study of User Communication Styles with Customer Service Agents versus Intelligent Virtual Agents

2020-07-01 · SIGDIAL (ACL) 2020 7 · Timothy Hewitt, Ian Beaver

We investigate differences in user communication with live chat agents versus a commercial Intelligent Virtual Agent (IVA). This case study compares the two types of interactions in the same domain for the same company f…

ChatbotDiversity

Quantity versus quality in publication activity: knowledge production at the regional level

2023-11-15 · Timur Gareev, Irina Peker

This study contributes to the ongoing debate regarding the balance between quality and quantity in research productivity and publication activity. Using empirical regional knowledge production functions, we establish a s…

Consistent Quantity-Quality Control across Scenes for Deployment-Aware Gaussian Splatting

2025-05-15 · Fengdi Zhang, Hongkun Cao, Ruqi Huang

To reduce storage and computational costs, 3D Gaussian splatting (3DGS) seeks to minimize the number of Gaussians used while preserving high rendering quality, introducing an inherent trade-off between Gaussian quantity …

3DGSModel CompressionNovel View Synthesis

Quality or Quantity: Toward a Unified Approach for Multi-organ Segmentation in Body CT

2022-03-03 · Fakrul Islam Tushar, Husam Nujaim, Wanyi Fu, Ehsan Abadi 외

Organ segmentation of medical images is a key step in virtual imaging trials. However, organ segmentation datasets are limited in terms of quality (because labels cover only a few organs) and quantity (since case numbers…

Organ SegmentationSegmentation

MoEdit: On Learning Quantity Perception for Multi-object Image Editing

2025-03-13 · CVPR 2025 1 · Yanfeng Li, Kahou Chan, Yue Sun, ChanTong Lam 외

Multi-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications.…

AttributeImage GenerationObjectStyle Transfer