Mr. Right: Multimodal Retrieval on Representation of ImaGe witH Text
Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate information from multiple modalities. In Information Retrieval, traditional retrieval tasks focus on the similarity between unimodal documents and queries, while image-text retrieval hypothesizes that most texts contain the scene context from images. This separation has ignored that real-world queries may involve text content, image captions, or both. To address this, we introduce Multimodal Retrieval on Representation of ImaGe witH Text (Mr. Right), a novel and comprehensive dataset for multimodal retrieval. We utilize the Wikipedia dataset with rich text-image examples and generate three types of text-based queries with different modality information: text-related, image-related, and mixed. To validate the effectiveness of our dataset, we provide a multimodal training paradigm and evaluate previous text retrieval and image retrieval frameworks. The results show that proposed multimodal retrieval can improve retrieval performance, but creating a well-unified document representation with texts and images is still a challenge. We hope Mr. Right allows us to broaden current retrieval systems better and contributes to accelerating the advancement of multimodal learning in the Information Retrieval.
Code (1)
Tasks
Image CaptioningImage RetrievalImage-text RetrievalInformation RetrievalRetrievalText RetrievalSimilar Papers 제목 키워드 기반
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
Multimodal retrieval systems struggle to resolve image-text queries against text-only corpora: the best vision-language encoder achieves only 27.6 nDCG@10 on MM-BRIGHT, underperforming strong text-only retrievers. We arg…
Reinforcement LearningText RetrievalSafeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment
As Retrieval-Augmented Generation (RAG) evolves into service-oriented platforms (Rag-as-a-Service) with shared knowledge bases, protecting the copyright of contributed data becomes essential. Existing watermarking method…
RAGRetrieval-augmented GenerationContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising
Contextual advertising serves ads that are aligned to the content that the user is viewing. The rapid growth of video content on social platforms and streaming services, along with privacy concerns, has increased the nee…
RetrievalText to Video RetrievalVideo RetrievalCross-Modality Sub-Image Retrieval using Contrastive Multimodal Image Representations
In tissue characterization and cancer diagnostics, multimodal imaging has emerged as a powerful technique. Thanks to computational advances, large datasets can be exploited to discover patterns in pathologies and improve…
Content-Based Image RetrievalImage RetrievalRetrievalMultimodal semantic retrieval for product search
Semantic retrieval (also known as dense retrieval) based on textual data has been extensively studied for both web search and product search application fields, where the relevance of a query and a potential target docum…
RetrievalSemantic Retrieval