paper-with-me

Papers

Multi-level Attention Fusion Network for Audio-visual Event Recognition

2021-06-12 · Mathilde Brousmiche, Jean Rouat, Stéphane Dupont

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion network (MAFnet), an architecture that can dynamically fuse visual and audio information for event recognition. Inspired by prior studies in neuroscience, we couple both modalities at different levels of visual and audio paths. Furthermore, the network dynamically highlights a modality at a given time window relevant to classify events. Experimental results in AVE (Audio-Visual Event), UCF51, and Kinetics-Sounds datasets show that the approach can effectively improve the accuracy in audio-visual event classification. Code is available at: https://github.com/numediart/MAFnet

📄 PDF Abstract BibTeX arXiv:2106.06736

Code (1)

numediart/MAFnet 공식 구현 tf

Similar Papers 제목 키워드 기반

Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention

2024-03-07 · R. Gnana Praveen, Jahangir Alam

Person or identity verification has been recently gaining a lot of attention using audio-visual fusion as faces and voices share close associations with each other. Conventional approaches based on audio-visual fusion re…

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

2026-05-02 · Mayesha Maliha R. Mithila, Mylene C. Q. Farias arxiv

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded wh…

MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

2024-01-07 · He Wang, Pengcheng Guo, Pan Zhou, Lei Xie

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with noise-invariant visual cues and improve…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learning+3

Dual-Path Cross-Modal Attention for better Audio-Visual Speech Extraction

2022-07-09 · Zhongweiyang Xu, Xulin Fan, Mark Hasegawa-Johnson

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and…

Speech ExtractionSpeech Separation

MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video Parsing

2021-11-24 · Jiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng 외

Recognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehe…

audio-visual event localizationVideo Understanding