paper-with-me

Papers

Cross-modal Contrastive Learning with Asymmetric Co-attention Network for Video Moment Retrieval

2023-12-12 · Love Panta, Prashant Shrestha, Brabeem Sapkota, Amrita Bhattarai, Suresh Manandhar, Anand Kumar Sah

Video moment retrieval is a challenging task requiring fine-grained interactions between video and text modalities. Recent work in image-text pretraining has demonstrated that most existing pretrained models suffer from information asymmetry due to the difference in length between visual and textual sequences. We question whether the same problem also exists in the video-text domain with an auxiliary need to preserve both spatial and temporal information. Thus, we evaluate a recently proposed solution involving the addition of an asymmetric co-attention network for video grounding tasks. Additionally, we incorporate momentum contrastive loss for robust, discriminative representation learning in both modalities. We note that the integration of these supplementary modules yields better performance compared to state-of-the-art models on the TACoS dataset and comparable results on ActivityNet Captions, all while utilizing significantly fewer parameters with respect to baseline.

📄 PDF Abstract BibTeX arXiv:2312.07435

Code (1)

love481/Cross-modal-Contrastive-Learning-with-Asymmetric-Co-attention-Network-for-Video-Moment-Retrieval 공식 구현 pytorch

Tasks

Contrastive LearningMoment RetrievalRepresentation LearningRetrievalVideo Grounding

Similar Papers 제목 키워드 기반

Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark

2026-08-26 · Bohan Deng, Shuo Ye, Zitong Yu arxiv

Audio-visual cross-modal Fine-Grained Visual Categorization (FGVC) aims to identify fine-grained categories by jointly leveraging visual and auditory information. However, FGVC under asymmetric cross-modal scenarios has …

Representation Learning

Rethinking Multimodal Content Moderation from an Asymmetric Angle with Mixed-modality

2023-05-17 · Jialin Yuan, Ye Yu, Gaurav Mittal, Matthew Hall 외

There is a rapidly growing need for multimodal content moderation (CM) as more and more content on social media is multimodal in nature. Existing unimodal CM systems may fail to catch harmful content that crosses modalit…

Exploiting Auxiliary Caption for Video Grounding

2023-01-15 · Hongxiang Li, Meng Cao, Xuxin Cheng, Zhihong Zhu 외

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the {sparsity dilemma} in video annotations, which fails to provide the context informa…

Contrastive LearningDense Video CaptioningSentenceVideo Captioning+1

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

2025-11-05 · Guozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng 외 arxiv

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we …

Video Generation

ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

2024-04-18 · Han Fang, Xianghao Zang, Chao Ban, Zerun Feng 외

Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content th…

DiversityRetrievalVideo Retrieval