paper-with-me

Papers

Leveraging Image-Text Similarity and Caption Modification for the DataComp Challenge: Filtering Track and BYOD Track

2023-10-23 · Shuhei Yokoo, Peifei Zhu, Yuchi Ishikawa, Mikihiro Tanaka, Masayoshi Kondo, Hirokatsu Kataoka

Large web crawl datasets have already played an important role in learning multimodal features with high generalization capabilities. However, there are still very limited studies investigating the details or improvements of data design. Recently, a DataComp challenge has been designed to propose the best training data with the fixed models. This paper presents our solution to both filtering track and BYOD track of the DataComp challenge. Our solution adopts large multimodal models CLIP and BLIP-2 to filter and modify web crawl data, and utilize external datasets along with a bag of tricks to improve the data quality. Experiments show our solution significantly outperforms DataComp baselines (filtering track: 6.6% improvement, BYOD track: 48.5% improvement).

📄 PDF Abstract BibTeX arXiv:2310.14581

Code (0)

등록된 구현이 없습니다.

Tasks

text similarity

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

CoLLM: A Large Language Model for Composed Image Retrieval

2025-03-25 · CVPR 2025 1 · Chuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah 외

Composed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modific…

Image RetrievalLanguage ModelingLanguage ModellingLarge Language Model+3

Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning

2024-02-21 · Antoine Chaffin, Ewa Kijak, Vincent Claveau

Training image captioning models using teacher forcing results in very generic samples, whereas more distinctive captions can be very useful in retrieval applications or to produce alternative texts describing images for…

Cross-Modal RetrievalImage CaptioningReinforcement Learning (RL)Retrieval

Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning

2024-09-23 · Siddharth Betala, Ishan Chokshi

In this paper, we describe our system under the team name Brotherhood for the English-to-Lowres Multi-Modal Translation Task. We participate in the multi-modal translation tasks for English-Hindi, English-Hausa, English-…

Image CaptioningSemantic SimilaritySemantic Textual SimilarityTranslation

Fine-Grained Zero-Shot Composed Image Retrieval with Complementary Visual-Semantic Integration

2026-01-20 · Yongcong Ye, Kai Zhang, Yanghai Zhang, Enhong Chen 외 arxiv

Zero-shot composed image retrieval (ZS-CIR) is a rapidly growing area with significant practical applications, allowing users to retrieve a target image by providing a reference image and a relative caption describing th…

Information ExtractionInformation RetrievalImage Retrieval

LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval

2024-07-11 · SIGIR 2024 7 · Zhenyu Yang, Dizhan Xue, Shengsheng Qian, WeiMing Dong 외

Zero-Shot Composed Image Retrieval (ZS-CIR) has garnered increasing interest in recent years, which aims to retrieve a target image based on a query composed of a reference image and a modification text without training …

Image RetrievalImage to textRetrievalZero-Shot Composed Image Retrieval (ZS-CIR)