paper-with-me

Papers

Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval

2025-04-03 · CVPR 2025 1 · Boseung Jeong, Jicheol Park, Sungyeon Kim, Suha Kwak

Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and textual features and often ignore audio, although it helps enhance overall comprehension of video content. Moreover, traditional models that incorporate audio blindly utilize the audio input regardless of whether it is useful or not, resulting in suboptimal video representation. To address these limitations, we propose a novel video-text retrieval framework, Audio-guided VIdeo representation learning with GATEd attention (AVIGATE), that effectively leverages audio cues through a gated attention mechanism that selectively filters out uninformative audio signals. In addition, we propose an adaptive margin-based contrastive loss to deal with the inherently unclear positive-negative relationship between video and text, which facilitates learning better video-text alignment. Our extensive experiments demonstrate that AVIGATE achieves state-of-the-art performance on all the public benchmarks.

📄 PDF Abstract BibTeX arXiv:2504.02397

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRepresentation LearningRetrievalText RetrievalVideo-Text RetrievalVideo Understanding

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation

2023-03-29 · Jiawei Liu, Weining Wang, Sihan Chen, Xinxin Zhu 외

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic vide…

Audio GenerationContrastive LearningDecoderVideo Generation

Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

2022-01-01 · CVPR 2022 1 · Wenwen Pan, Haonan Shi, Zhou Zhao, Jieming Zhu 외

Audio-Guided video semantic segmentation is a challenging problem in visual analysis and editing, which automatically separates foreground objects from background in a video sequence according to the referring audio …

DecoderDenoisingSegmentationSemantic Segmentation+2

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

2026-05-18 · Haojie Zheng, Yixin Yang, Siqi Yang, Shuchen Weng 외 arxiv

Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, …

Instruction FollowingVideo Generation

Temporal Cue Guided Video Highlight Detection With Low-Rank Audio-Visual Fusion

2021-01-01 · ICCV 2021 10 · Qinghao Ye, Xiyue Shen, Yuan Gao, ZiRui Wang 외

Video highlight detection plays an increasingly important role in social media content filtering, however, it remains highly challenging to develop automated video highlight detection methods because of the lack of t…

Highlight DetectionModel Optimization

Past and Future Motion Guided Network for Audio Visual Event Localization

2022-05-08 · Tingxiu Chen, Jianqin Yin, Jin Tang

In recent years, audio-visual event localization has attracted much attention. It's purpose is to detect the segment containing audio-visual events and recognize the event category from untrimmed videos. Existing methods…

audio-visual event localization