paper-with-me

홈 › Papers

Multimodal Information Retrieval for Open World with Edit Distance Weak Supervision

2025-06-25 · KMA Solaiman, Bharat Bhargava

Existing multi-media retrieval models either rely on creating a common subspace with modality-specific representation models or require schema mapping among modalities to measure similarities among multi-media data. Our goal is to avoid the annotation overhead incurred from considering retrieval as a supervised classification task and re-use the pretrained encoders in large language models and vision tasks. We propose "FemmIR", a framework to retrieve multimodal results relevant to information needs expressed with multimodal queries by example without any similarity label. Such identification is necessary for real-world applications where data annotations are scarce and satisfactory performance is required without fine-tuning with a common framework across applications. We curate a new dataset called MuQNOL for benchmarking progress on this task. Our technique is based on weak supervision introduced through edit distance between samples: graph edit distance can be modified to consider the cost of replacing a data sample in terms of its properties, and relevance can be measured through the implicit signal from the amount of edit cost among the objects. Unlike metric learning or encoding networks, FemmIR re-uses the high-level properties and maintains the property value and relationship constraints with a multi-level interaction score between data samples and the query example provided by the user. We empirically evaluate FemmIR on a missing person use case with MuQNOL. FemmIR performs comparably to similar retrieval systems in delivering on-demand retrieval results with exact and approximate similarities while using the existing property identifiers in the system.

📄 PDF Abstract BibTeX arXiv:2506.20070

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingInformation RetrievalMetric LearningRetrieval

Similar Papers 제목 키워드 기반

A Survey of Multimodal Composite Editing and Retrieval

2024-09-09 · Suyan Li, Fuxiang Huang, Lei Zhang

In the real world, where information is abundant and diverse across different modalities, understanding and utilizing various data types to improve retrieval systems is a key focus of research. Multimodal composite retri…

RetrievalSurvey

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

2026-08-15 · Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun 외 hf

Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how mult…

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

2026-07-09 · Feng Wang, Canmiao Fu, Zhipeng Huang, Chen Li 외 arxiv

Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a sh…

Reinforcement LearningImage Generation

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

2026-01-30 · Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai 외 arxiv

In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open-domain video shot retrieval, which inv…

Information RetrievalVideo Retrieval

MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

2024-10-15 · CVPR 2025 1 · Reno Kriz, Kate Sanders, David Etter, Kenton Murray 외

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on…

DescriptiveRetrievalVideo Retrieval