paper-with-me

Papers

MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed

2025-06-11 · Jiaqi Samantha Zhan, Crystina Zhang, Shengyao Zhuang, Xueguang Ma, Jimmy Lin

Effective video retrieval remains challenging due to the complexity of integrating visual, auditory, and textual modalities. In this paper, we explore unified retrieval methods using OmniEmbed, a powerful multimodal embedding model from the Tevatron 2.0 toolkit, in the context of the MAGMaR shared task. Evaluated on the comprehensive MultiVENT 2.0 dataset, OmniEmbed generates unified embeddings for text, images, audio, and video, enabling robust multimodal retrieval. By finetuning OmniEmbed with the combined multimodal data--visual frames, audio tracks, and textual descriptions provided in MultiVENT 2.0, we achieve substantial improvements in complex, multilingual video retrieval tasks. Our submission achieved the highest score on the MAGMaR shared task leaderboard among public submissions as of May 20th, 2025, highlighting the practical effectiveness of our unified multimodal retrieval approach. Model checkpoint in this work is opensourced.

📄 PDF Abstract BibTeX arXiv:2506.09409

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

Findings of the MAGMaR 2026 Shared Task

2026-06-10 · Alexander Martin, Dengjia Zhang, Joel Brogan, Francis Ferraro 외 arxiv

This overview paper presents the results of the shared task for the second workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR). In this shared task participants submitted systems focused on eith…

Video Retrieval

CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

2026-05-18 · Mahesh Bhosale, Abdul Wasi, Vishvesh Trivedi, Pengyu Yan 외 arxiv

Grounded multi-video question answering over real-world news events requires systems to surface query-relevant evidence across heterogeneous video archives while attributing every claim to its supporting source. We intro…

Video Question Answering

MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation

2026-05-17 · Debashish Chakraborty, Dengjia Zhang, Jialiang Jin, Hanting Liu 외 arxiv

Retrieval-augmented generation from videos requires systems to retrieve relevant audiovisual evidence from large corpora and synthesize it into coherent, attributed text. Current approaches struggle at both ends: retriev…

Video Retrieval

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation

2026-05-16 · Pengyu Yan, Akhil Gorugantu, Mahesh Bhosale, Abdul Wasi 외 arxiv

Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora. Existing large vision-language models (LVLMs) often underperform in…

Object DetectionVisual Reasoning

Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

2026-06-06 · Jiaxin Dai, Zehang Wei, Jiamin Yan, Xiang Xiang arxiv

This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via MultimodAl Retrieval (MAGMaR). Addressing the critical challenges of cross-lingual long-video comprehension, strict p…

Information RetrievalSemantic RetrievalLogical ReasoningVideo Retrieval