paper-with-me

홈 › Papers

C-CLIP: Contrastive Image-Text Encoders to Close the Descriptive-Commentative Gap

2023-09-06 · William Theisen, Walter Scheirer

The interplay between the image and comment on a social media post is one of high importance for understanding its overall message. Recent strides in multimodal embedding models, namely CLIP, have provided an avenue forward in relating image and text. However the current training regime for CLIP models is insufficient for matching content found on social media, regardless of site or language. Current CLIP training data is based on what we call `descriptive'' text: text in which an image is merely described. This is something rarely seen on social media, where the vast majority of text content is `commentative'' in nature. The captions provide commentary and broader context related to the image, rather than describing what is in it. Current CLIP models perform poorly on retrieval tasks where image-caption pairs display a commentative relationship. Closing this gap would be beneficial for several important application areas related to social media. For instance, it would allow groups focused on Open-Source Intelligence Operations (OSINT) to further aid efforts during disaster events, such as the ongoing Russian invasion of Ukraine, by easily exposing data to non-technical users for discovery and analysis. In order to close this gap we demonstrate that training contrastive image-text encoders on explicitly commentative pairs results in large improvements in retrieval results, with the results extending across a variety of non-English languages.

📄 PDF Abstract BibTeX arXiv:2309.03921

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveRetrieval

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge

2024-10-16 · Fawaz Sammani, Nikos Deligiannis

Contrastive Language-Image Pretraining (CLIP) performs zero-shot image classification by mapping images and textual class representation into a shared embedding space, then retrieving the class closest to the image. This…

Classificationimage-classificationImage Classificationzero-shot-classification+2

Enhancing CLIP Conceptual Embedding through Knowledge Distillation

2024-12-04 · Kuei-Chun Kao

Recently, CLIP has become an important model for aligning images and text in multi-modal contexts. However, researchers have identified limitations in the ability of CLIP's text and image encoders to extract detailed kno…

Contrastive LearningKnowledge Distillation

On the Difference of BERT-style and CLIP-style Text Encoders

2023-06-06 · Zhihong Chen, Guiming Hardy Chen, Shizhe Diao, Xiang Wan 외

Masked language modeling (MLM) has been one of the most popular pretraining recipes in natural language processing, e.g., BERT, one of the representative models. Recently, contrastive language-image pretraining (CLIP) ha…

Image GenerationLanguage ModelingLanguage ModellingMasked Language Modeling+2

Expediting Contrastive Language-Image Pretraining via Self-distilled Encoders

2023-12-19 · Bumsoo Kim, Jinhyung Kim, Yeonsik Jo, Seung Hwan Kim

Recent advances in vision language pretraining (VLP) have been largely attributed to the large-scale data collected from the web. However, uncurated dataset contains weakly correlated image-text pairs, causing data ineff…

Knowledge Distillation

Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning

2022-12-09 · CVPR 2023 1 · Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang 외

We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of the vision encoder and the CLS token of t…

Contrastive Learningimage-classificationImage ClassificationOpen Vocabulary Semantic Segmentation+6