paper-with-me

Papers

Audio-Visual Event Localization via Recursive Fusion by Joint Co-Attention

2020-08-14 · Bin Duan, Hao Tang, Wei Wang, Ziliang Zong, Guowei Yang, Yan Yan

The major challenge in audio-visual event localization task lies in how to fuse information from multiple modalities effectively. Recent works have shown that attention mechanism is beneficial to the fusion process. In this paper, we propose a novel joint attention mechanism with multimodal fusion methods for audio-visual event localization. Particularly, we present a concise yet valid architecture that effectively learns representations from multiple modalities in a joint manner. Initially, visual features are combined with auditory features and then turned into joint representations. Next, we make use of the joint representations to attend to visual features and auditory features, respectively. With the help of this joint co-attention, new visual and auditory features are produced, and thus both features can enjoy the mutually improved benefits from each other. It is worth noting that the joint co-attention unit is recursive meaning that it can be performed multiple times for obtaining better joint representations progressively. Extensive experiments on the public AVE dataset have shown that the proposed method achieves significantly better results than the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2008.06581

Code (0)

등록된 구현이 없습니다.

Tasks

audio-visual event localizationvalid

Similar Papers 제목 키워드 기반

Audio-Visual Event Localization in Unconstrained Videos

2018-03-23 · ECCV 2018 9 · Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan 외

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio…

audio-visual event localizationTemporal Localization

Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios

2024-06-21 · Ya Jiang, Qing Wang, Jun Du, Maocheng Hu 외

This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal lea…

Data AugmentationSound Event Localization and Detection

Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection

2023-12-14 · Davide Berghi, Peipei Wu, Jinzheng Zhao, Wenwu Wang 외

Sound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has bee…

Data AugmentationEvent DetectionSound Event DetectionSound Event Localization and Detection

Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source Localization

2023-08-11 · Sung Jin Um, DongJin Kim, Jung Uk Kim

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, e…

Sound Source Localization

MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video Parsing

2021-11-24 · Jiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng 외

Recognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehe…

audio-visual event localizationVideo Understanding