paper-with-me

Papers

Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning

2022-07-15 · Shibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang, Tatsunori Hashimoto

The development of CLIP [Radford et al., 2021] has sparked a debate on whether language supervision can result in vision models with more transferable representations than traditional image-only methods. Our work studies this question through a carefully controlled comparison of two approaches in terms of their ability to learn representations that generalize to downstream classification tasks. We find that when the pre-training dataset meets certain criteria -- it is sufficiently large and contains descriptive captions with low variability -- image-only methods do not match CLIP's transfer performance, even when they are trained with more image data. However, contrary to what one might expect, there are practical settings in which these criteria are not met, wherein added supervision through captions is actually detrimental. Motivated by our findings, we devise simple prescriptions to enable CLIP to better leverage the language information present in existing pre-training datasets.

📄 PDF Abstract BibTeX arXiv:2207.07635

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveRepresentation Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

A Picture is Worth a Thousand Words: A Unified System for Diverse Captions and Rich Images Generation

2021-10-19 · Yupan Huang, Bei Liu, Jianlong Fu, Yutong Lu

A creative image-and-text generative AI system mimics humans' extraordinary abilities to provide users with diverse and comprehensive caption suggestions, as well as rich image creations. In this work, we demonstrate suc…

Diversity

Aesthetic Critiques Generation for Photos

2017-10-01 · ICCV 2017 10 · Kuang-Yu Chang, Kung-Hung Lu, Chu-Song Chen

It is said that a picture is worth a thousand words. Thus, there are various ways to describe an image, especially in aesthetic quality analysis. Although aesthetic quality assessment has generated a great deal of intere…

Image Captioning

ClipMatrix: Text-controlled Creation of 3D Textured Meshes

2021-09-27 · Nikolay Jetchev

If a picture is worth thousand words, a moving 3d shape must be worth a million. We build upon the success of recent generative methods that create images fitting the semantics of a text prompt, and extend it to the cont…

How Culturally Aware are Vision-Language Models?

2024-05-24 · Olena Burda-Lassen, Aman Chadha, Shashank Goswami, Vinija Jain

An image is often said to be worth a thousand words, and certain images can tell rich and insightful stories. Can these stories be told via image captioning? Images from folklore genres, such as mythology, folk dance, cu…

Image Captioning

A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation

2023-10-25 · Eyal Segalis, Dani Valevski, Danny Lumen, Yossi Matias 외

Text-to-image diffusion models achieved a remarkable leap in capabilities over the last few years, enabling high-quality and diverse synthesis of images from a textual prompt. However, even the most advanced models often…

Image CaptioningImage Generation