paper-with-me

홈 › Papers

Unified Coarse-to-Fine Alignment for Video-Text Retrieval

2023-09-18 · ICCV 2023 1 · Ziyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius, Mohit Bansal

The canonical approach to video-text retrieval leverages a coarse-grained or fine-grained alignment between visual and textual information. However, retrieving the correct video according to the text query is often challenging as it requires the ability to reason about both high-level (scene) and low-level (object) visual clues and how they relate to the text query. To this end, we propose a Unified Coarse-to-fine Alignment model, dubbed UCoFiA. Specifically, our model captures the cross-modal similarity information at different granularity levels. To alleviate the effect of irrelevant visual clues, we also apply an Interactive Similarity Aggregation module (ISA) to consider the importance of different visual features while aggregating the cross-modal similarity to obtain a similarity score for each granularity. Finally, we apply the Sinkhorn-Knopp algorithm to normalize the similarities of each level before summing them, alleviating over- and under-representation issues at different levels. By jointly considering the crossmodal similarity of different granularity, UCoFiA allows the effective unification of multi-grained alignments. Empirically, UCoFiA outperforms previous state-of-the-art CLIP-based methods on multiple video-text retrieval benchmarks, achieving 2.4%, 1.4% and 1.3% improvements in text-to-video retrieval R@1 on MSR-VTT, Activity-Net, and DiDeMo, respectively. Our code is publicly available at https://github.com/Ziyang412/UCoFiA.

📄 PDF Abstract BibTeX arXiv:2309.10091

Code (1)

ziyang412/ucofia 공식 구현 pytorch

Tasks

RetrievalText RetrievalText to Video RetrievalVideo RetrievalVideo-Text Retrieval

Similar Papers 제목 키워드 기반

Error-Propagation-Free Learned Video Compression With Dual-Domain Progressive Temporal Alignment

2025-12-11 · Han Li, Shaohui Li, Wenrui Dai, Chenglin Li 외 arxiv

Existing frameworks for learned video compression suffer from a dilemma between inaccurate temporal alignment and error propagation for motion estimation and compensation (ME/MC). The separate-transform framework employs…

TokenBinder: Text-Video Retrieval with One-to-Many Alignment Paradigm

2024-09-30 · Bingqing Zhang, Zhuo Cao, Heming Du, Xin Yu 외

Text-Video Retrieval (TVR) methods typically match query-candidate pairs by aligning text and video features in coarse-grained, fine-grained, or combined (coarse-to-fine) manners. However, these frameworks predominantly …

RetrievalVideo Retrieval

Multi-granularity Correspondence Learning from Long-term Noisy Videos

2024-01-30 · Yijie Lin, Jie Zhang, Zhenyu Huang, Jia Liu 외

Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling long videos. To address this issue, one …

Action SegmentationLong Video Retrieval (Background Removed)Video RetrievalVideo Understanding

Coarse to Fine: Video Retrieval before Moment Localization

2021-10-14 · Zijian Gao, Huanyu Liu, Jingyu Liu

The current state-of-the-art methods for video corpus moment retrieval (VCMR) often use similarity-based feature alignment approach for the sake of convenience and speed. However, late fusion methods like cosine similari…

Moment RetrievalRetrievalVideo Corpus Moment RetrievalVideo Retrieval

Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment

2024-06-27 · Hao Fei, Shengqiong Wu, Meishan Zhang, Min Zhang 외

While pre-training large-scale video-language models (VLMs) has shown remarkable potential for various downstream video-language tasks, existing VLMs can still suffer from certain commonly seen limitations, e.g., coarse-…