paper-with-me

홈 › Papers

Mr. Right: Multimodal Retrieval on Representation of ImaGe witH Text

2022-09-28 · Cheng-An Hsieh, Cheng-Ping Hsieh, Pu-Jen Cheng

Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate information from multiple modalities. In Information Retrieval, traditional retrieval tasks focus on the similarity between unimodal documents and queries, while image-text retrieval hypothesizes that most texts contain the scene context from images. This separation has ignored that real-world queries may involve text content, image captions, or both. To address this, we introduce Multimodal Retrieval on Representation of ImaGe witH Text (Mr. Right), a novel and comprehensive dataset for multimodal retrieval. We utilize the Wikipedia dataset with rich text-image examples and generate three types of text-based queries with different modality information: text-related, image-related, and mixed. To validate the effectiveness of our dataset, we provide a multimodal training paradigm and evaluate previous text retrieval and image retrieval frameworks. The results show that proposed multimodal retrieval can improve retrieval performance, but creating a well-unified document representation with texts and images is still a challenge. We hope Mr. Right allows us to broaden current retrieval systems better and contributes to accelerating the advancement of multimodal learning in the Information Retrieval.

📄 PDF Abstract BibTeX arXiv:2209.13764

Code (1)

hsiehjackson/mr.right 공식 구현 pytorch

Tasks

Image CaptioningImage RetrievalImage-text RetrievalInformation RetrievalRetrievalText Retrieval

Similar Papers 제목 키워드 기반

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment

2026-04-08 · Mohamed Darwish Mounis, Mohamed Mahmoud, Shaimaa Sedek, Mahmoud Abdalla 외 arxiv

Multimodal retrieval systems struggle to resolve image-text queries against text-only corpora: the best vision-language encoder achieves only 27.6 nDCG@10 on MM-BRIGHT, underperforming strong text-only retrievers. We arg…

Reinforcement LearningText Retrieval

Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment

2025-06-10 · Tianyu Chen, Jian Lou, Wenjie Wang

As Retrieval-Augmented Generation (RAG) evolves into service-oriented platforms (Rag-as-a-Service) with shared knowledge bases, protecting the copyright of contributed data becomes essential. Existing watermarking method…

RAGRetrieval-augmented Generation

ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising

2024-10-29 · Ashutosh Chaubey, Anoubhav Agarwaal, Sartaki Sinha Roy, Aayush Agrawal 외

Contextual advertising serves ads that are aligned to the content that the user is viewing. The rapid growth of video content on social platforms and streaming services, along with privacy concerns, has increased the nee…

RetrievalText to Video RetrievalVideo Retrieval

Cross-Modality Sub-Image Retrieval using Contrastive Multimodal Image Representations

2022-01-10 · Eva Breznik, Elisabeth Wetzer, Joakim Lindblad, Nataša Sladoje

In tissue characterization and cancer diagnostics, multimodal imaging has emerged as a powerful technique. Thanks to computational advances, large datasets can be exploited to discover patterns in pathologies and improve…

Content-Based Image RetrievalImage RetrievalRetrieval

Multimodal semantic retrieval for product search

2025-01-13 · Dong Liu, Esther Lopez Ramos

Semantic retrieval (also known as dense retrieval) based on textual data has been extensively studied for both web search and product search application fields, where the relevance of a query and a potential target docum…

RetrievalSemantic Retrieval