paper-with-me

홈 › Papers

Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP

2024-09-23 · Zeliang Zhang, Zhuo Liu, Mingqian Feng, Chenliang Xu

CLIP has demonstrated great versatility in adapting to various downstream tasks, such as image editing and generation, visual question answering, and video understanding. However, CLIP-based applications often suffer from misunderstandings regarding user intent, leading to discrepancies between the required number of objects and the actual outputs in image generation tasks. In this work, we empirically investigate the quantity bias in CLIP. By carefully designing different experimental settings and datasets, we comprehensively evaluate CLIP's understanding of quantity from text, image, and cross-modal perspectives. Our experimental results reveal a quantity bias in CLIP embeddings, impacting the reliability of downstream tasks.

📄 PDF Abstract BibTeX arXiv:2409.15035

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationQuestion AnsweringVideo UnderstandingVisual Question Answering

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Automatic classification of eclipsing binary stars using deep learning methods

2021-08-03 · Michal Čokina, Viera Maslej-Krešňáková, Peter Butka, Štefan Parimucha

In the last couple of decades, tremendous progress has been achieved in developing robotic telescopes and, as a result, sky surveys (both terrestrial and space) have become the source of a substantial amount of new obser…

ClassificationDeep Learning

The TESS Ten Thousand Catalog: 10,001 uniformly-vetted and -validated Eclipsing Binary Stars detected in Full-Frame Image data by machine learning and analyzed by citizen scientists

2025-06-05 · Veselin B. Kostov, Brian P. Powell, Aline U. Fornear, Marco Z. Di Fraia 외

The Transiting Exoplanet Survey Satellite (TESS) has surveyed nearly the entire sky in Full-Frame Image mode with a time resolution of 200 seconds to 30 minutes and a temporal baseline of at least 27 days. In addition to…

Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP

2022-08-10 · Thao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh 외

Web-crawled datasets have enabled remarkable generalization capabilities in recent image-text models such as CLIP (Contrastive Language-Image pre-training) or Flamingo, but little is known about the dataset creation proc…

Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency

2025-09-29 · Jiaqi Tan, Fangyu Li, Yang Liu arxiv

Instruction driven image editing with standard CLIP text encoders often fails in complex scenes with many objects. We present QL-Adapter, a framework for multiple object editing that tackles two challenges: enforcing obj…

Instruction FollowingImage Editing

Cataloging Accreted Stars within Gaia DR2 using Deep Learning

2019-07-15 · Bryan Ostdiek, Lina Necib, Timothy Cohen, Marat Freytsis 외

The goal of this study is to present the development of a machine learning based approach that utilizes phase space alone to separate the Gaia DR2 stars into two categories: those accreted onto the Milky Way from those t…

Deep LearningTransfer Learning