paper-with-me

홈 › Papers

X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval

2022-03-28 · CVPR 2022 1 · Satya Krishna Gorti, Noel Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, Guangwei Yu

In text-video retrieval, the objective is to learn a cross-modal similarity function between a text and a video that ranks relevant text-video pairs higher than irrelevant pairs. However, videos inherently express a much wider gamut of information than texts. Instead, texts often capture sub-regions of entire videos and are most semantically similar to certain frames within videos. Therefore, for a given text, a retrieval model should focus on the text's most semantically similar video sub-regions to make a more relevant comparison. Yet, most existing works aggregate entire videos without directly considering text. Common text-agnostic aggregations schemes include mean-pooling or self-attention over the frames, but these are likely to encode misleading visual information not described in the given text. To address this, we propose a cross-modal attention model called X-Pool that reasons between a text and the frames of a video. Our core mechanism is a scaled dot product attention for a text to attend to its most semantically similar frames. We then generate an aggregated video representation conditioned on the text's attention weights over the frames. We evaluate our method on three benchmark datasets of MSR-VTT, MSVD and LSMDC, achieving new state-of-the-art results by up to 12% in relative improvement in Recall@1. Our findings thereby highlight the importance of joint text-video reasoning to extract important visual cues according to text. Full code and demo can be found at: https://layer6ai-labs.github.io/xpool/

📄 PDF Abstract BibTeX arXiv:2203.15086

Code (1)

layer6ai-labs/xpool 공식 구현 pytorch

Tasks

RetrievalText to Video RetrievalVideo RetrievalVideo-Text Retrieval

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
EXP-$Does Expedia refund a cancelled flight? EXP-$Does Expedia refund a cancelled flight? If you’re wondering, +1 888-829-0881 does Expedia refund a cancelled flight, the answer depends+1 888-829-0881 on the airline’s…

Similar Papers 제목 키워드 기반

CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms

2025-05-22 · Shilin Yan, Jiaming Han, Joey Tsai, Hongwei Xue 외

The advent of Large Multimodal Models (LMMs) has significantly enhanced Large Language Models (LLMs) to process and interpret diverse data modalities (e.g., image and video). However, as input complexity increases, parti…

Token Reduction

VLG-Net: Video-Language Graph Matching Network for Video Grounding

2020-11-19 · Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner 외

Grounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query. The solution to this challenging task demands understanding videos' and queries' semantic …

Graph MatchingMoment RetrievalNatural Language Moment RetrievalTemporal Localization+1

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

2025-10-24 · Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer 외 arxiv

Foley Control is a lightweight approach to video-guided Foley that keeps pretrained single-modality models frozen and learns only a small cross-attention bridge between them. We connect V-JEPA2 video embeddings to a froz…

SBAT: Video Captioning with Sparse Boundary-Aware Transformer

2020-07-23 · Tao Jin, Siyu Huang, Ming Chen, Yingming Li 외

In this paper, we focus on the problem of applying the transformer structure to video captioning effectively. The vanilla transformer is proposed for uni-modal language generation task such as machine translation. Howeve…

Machine Translationmultimodal interactionText GenerationTranslation+1

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors

2026-07-17 · Yilin Wang, Xiangxi Zheng, Dongxing Mao, Linjie Li 외 arxiv

Understanding long videos with multimodal large language models (MLLMs) requires selecting a compact set of frames from thousands of candidates, yet identifying the right frames seemingly requires understanding the video…