paper-with-me

홈 › Papers

MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline

2024-07-17 · Donghoon Han, Eunhwan Park, Gisang Lee, Adam Lee, Nojun Kwak

The rapid expansion of multimedia content has made accurately retrieving relevant videos from large collections increasingly challenging. Recent advancements in text-video retrieval have focused on cross-modal interactions, large-scale foundation model training, and probabilistic modeling, yet often neglect the crucial user perspective, leading to discrepancies between user queries and the content retrieved. To address this, we introduce MERLIN (Multimodal Embedding Refinement via LLM-based Iterative Navigation), a novel, training-free pipeline that leverages Large Language Models (LLMs) for iterative feedback learning. MERLIN refines query embeddings from a user perspective, enhancing alignment between queries and video content through a dynamic question answering process. Experimental results on datasets like MSR-VTT, MSVD, and ActivityNet demonstrate that MERLIN substantially improves Recall@1, outperforming existing systems and confirming the benefits of integrating LLMs into multimodal retrieval systems for more responsive and context-aware multimedia retrieval.

📄 PDF Abstract BibTeX arXiv:2407.12508

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringRetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

Merlin HugeCTR: GPU-accelerated Recommender System Training and Inference

2022-10-17 · Joey Wang, Yingcan Wei, Minseok Lee, Matthias Langer 외

In this talk, we introduce Merlin HugeCTR. Merlin HugeCTR is an open source, GPU-accelerated integration framework for click-through rate estimation. It optimizes both training and inference, whilst enabling model traini…

CPUGPURecommendation SystemsRetrieval

MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking

2025-10-16 · Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja, Zaid Sheikh 외 arxiv

This paper introduces MERLIN, a novel testbed system for the task of Multilingual Multimodal Entity Linking. The created dataset includes BBC news article titles, paired with corresponding images, in five languages: Hind…

Entity Linking

DiaLoc: An Iterative Approach to Embodied Dialog Localization

2024-03-11 · CVPR 2024 1 · Chao Zhang, Mohan Li, Ignas Budvytis, Stephan Liwicki

Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing…

Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

2026-09-09 · Parinthapat Pengpun, Simran Khanuja, Graham Neubig hf

Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pa…

Entity Linking

Show me your NFT and I tell you how it will perform: Multimodal representation learning for NFT selling price prediction

2023-02-03 · Davide Costa, Lucio La Cava, Andrea Tagarelli

Non-Fungible Tokens (NFTs) represent deeds of ownership, based on blockchain technologies and smart contracts, of unique crypto assets on digital art forms (e.g., artworks or collectibles). In the spotlight after skyrock…

Graph Neural NetworkMultimodal Deep LearningRepresentation Learning