paper-with-me

Papers

Dual Attention Matching for Audio-Visual Event Localization

2019-10-01 · ICCV 2019 10 · Yu Wu, Linchao Zhu, Yan Yan, Yi Yang

In this paper, we investigate the audio-visual event localization problem. This task is to localize a visible and audible event in a video. Previous methods first divide a video into short segments, and then fuse visual and acoustic features at the segment level. The duration of these segments is usually short, making the visual and acoustic feature of each segment possibly not well aligned. Direct concatenation of the two features at the segment level can be vulnerable to a minor temporal misalignment of the two signals. We propose a Dual Attention Matching (DAM) module to cover a longer video duration for better high-level event information modeling, while the local temporal information is attained by the global cross-check mechanism. Our premise is that one should watch the whole video to understand the high-level event, while shorter segments should be checked in detail for localization. Specifically, the global feature of one modality queries the local feature in the other modality in a bi-directional way. With temporal co-occurrence encoded between auditory and visual signals, DAM can be readily applied in various audio-visual event localization tasks, e.g., cross-modality localization, supervised event localization. Experiments on the AVE dataset show our method outperforms the state-of-the-art by a large margin.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

audio-visual event localization

Similar Papers 제목 키워드 기반

Audio-Visual Event Localization in Unconstrained Videos

2018-03-23 · ECCV 2018 9 · Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan 외

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio…

audio-visual event localizationTemporal Localization

DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching

2025-08-08 · Wei Chen, Binzhu Sha, Dan Luo, Jing Yang 외 arxiv

Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degrad…

Self-Supervised LearningAudio GenerationVoice Conversion

Past and Future Motion Guided Network for Audio Visual Event Localization

2022-05-08 · Tingxiu Chen, Jianqin Yin, Jin Tang

In recent years, audio-visual event localization has attracted much attention. It's purpose is to detect the segment containing audio-visual events and recognize the event category from untrimmed videos. Existing methods…

audio-visual event localization

Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video Parsing

2021-06-19 · CVPR 2021 1 · Yu Wu, Yi Yang

We investigate the weakly-supervised audio-visual video parsing task, which aims to parse a video into temporal event segments and predict the audible or visible event categories. The task is challenging since there …

Contrastive Learning

Label-anticipated Event Disentanglement for Audio-Visual Video Parsing

2024-07-11 · Jinxing Zhou, Dan Guo, Yuxin Mao, Yiran Zhong 외

Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identification challenging. While traditional met…

audio-visual event localizationDisentanglementSemantic SimilaritySemantic Textual Similarity