paper-with-me

Papers

Multi-Modulation Network for Audio-Visual Event Localization

2021-08-26 · Hao Wang, Zheng-Jun Zha, Liang Li, Xuejin Chen, Jiebo Luo

We study the problem of localizing audio-visual events that are both audible and visible in a video. Existing works focus on encoding and aligning audio and visual features at the segment level while neglecting informative correlation between segments of the two modalities and between multi-scale event proposals. We propose a novel MultiModulation Network (M2N) to learn the above correlation and leverage it as semantic guidance to modulate the related auditory, visual, and fused features. In particular, during feature encoding, we propose cross-modal normalization and intra-modal normalization. The former modulates the features of two modalities by establishing and exploiting the cross-modal relationship. The latter modulates the features of a single modality with the event-relevant semantic guidance of the same modality. In the fusion stage,we propose a multi-scale proposal modulating module and a multi-alignment segment modulating module to introduce multi-scale event proposals and enable dense matching between cross-modal segments. With the auditory, visual, and fused features modulated by the correlation information regarding audio-visual events, M2N performs accurate event localization. Extensive experiments conducted on the AVE dataset demonstrate that our proposed method outperforms the state of the art in both supervised event localization and cross-modality localization.

📄 PDF Abstract BibTeX arXiv:2108.11773

Code (0)

등록된 구현이 없습니다.

Tasks

audio-visual event localization

Similar Papers 제목 키워드 기반

Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization

2024-09-12 · Ling Xing, Hongyu Qu, Rui Yan, Xiangbo Shu 외

Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying dura…

cross-modal alignment

Audio-Visual Event Localization in Unconstrained Videos

2018-03-23 · ECCV 2018 9 · Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan 외

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio…

audio-visual event localizationTemporal Localization

MPN: Multimodal Parallel Network for Audio-Visual Event Localization

2021-04-07 · Jiashuo Yu, Ying Cheng, Rui Feng

Audio-visual event localization aims to localize an event that is both audible and visible in the wild, which is a widespread audio-visual scene analysis task for unconstrained videos. To address this task, we propose a …

audio-visual event localizationGeneral Classification

CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization

2024-08-04 · Xiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao 외

The audio-visual event localization task requires identifying concurrent visual and auditory events from unconstrained videos within a network model, locating them, and classifying their category. The efficient extractio…

audio-visual event localization

AVECL-UMONS database for audio-visual event classification and localization

2020-10-02 · Mathilde Brousmiche, Stéphane Dupont, Jean Rouat

We introduce the AVECL-UMons dataset for audio-visual event classification and localization in the context of office environments. The audio-visual dataset is composed of 11 event classes recorded at several realistic po…

General Classification