paper-with-me

홈 › Papers

Cooperative Dual Attention for Audio-Visual Speech Enhancement with Facial Cues

2023-11-24 · Feixiang Wang, Shuang Yang, Shiguang Shan, Xilin Chen

In this work, we focus on leveraging facial cues beyond the lip region for robust Audio-Visual Speech Enhancement (AVSE). The facial region, encompassing the lip region, reflects additional speech-related attributes such as gender, skin color, nationality, etc., which contribute to the effectiveness of AVSE. However, static and dynamic speech-unrelated attributes also exist, causing appearance changes during speech. To address these challenges, we propose a Dual Attention Cooperative Framework, DualAVSE, to ignore speech-unrelated information, capture speech-related information with facial cues, and dynamically integrate it with the audio signal for AVSE. Specifically, we introduce a spatial attention-based visual encoder to capture and enhance visual speech information beyond the lip region, incorporating global facial context and automatically ignoring speech-unrelated information for robust visual feature extraction. Additionally, a dynamic visual feature fusion strategy is introduced by integrating a temporal-dimensional self-attention module, enabling the model to robustly handle facial variations. The acoustic noise in the speaking process is variable, impacting audio quality. Therefore, a dynamic fusion strategy for both audio and visual features is introduced to address this issue. By integrating cooperative dual attention in the visual encoder and audio-visual fusion strategy, our model effectively extracts beneficial speech information from both audio and visual cues for AVSE. Thorough analysis and comparison on different datasets, including normal and challenging cases with unreliable or absent visual information, consistently show our model outperforming existing methods across multiple metrics.

📄 PDF Abstract BibTeX arXiv:2311.14275

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention

2024-04-21 · Xinxin Jiao, Liejun Wang, Yinfeng Yu

Speech emotion recognition is crucial in human-computer interaction, but extracting and using emotional cues from audio poses challenges. This paper introduces MFHCA, a novel method for Speech Emotion Recognition using M…

Emotion RecognitionSpeech Emotion Recognition

Dual-Path Cross-Modal Attention for better Audio-Visual Speech Extraction

2022-07-09 · Zhongweiyang Xu, Xulin Fan, Mark Hasegawa-Johnson

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and…

Speech ExtractionSpeech Separation

Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture

2018-09-28 · Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Georgios Tzimiropoulos 외

Recent works in speech recognition rely either on connectionist temporal classification (CTC) or sequence-to-sequence models for character-level recognition. CTC assumes conditional independence of individual characters,…

Audio-Visual Speech RecognitionAutomatic Speech Recognition (ASR)Lipreadingspeech-recognition+2

An Empirical Analysis of Deep Audio-Visual Models for Speech Recognition

2018-12-21 · Devesh Walawalkar, Yihui He, Rohit Pillai

In this project, we worked on speech recognition, specifically predicting individual words based on both the video frames and audio. Empowered by convolutional neural networks, the recent speech recognition and lip readi…

Lip ReadingSensitivityspeech-recognitionSpeech Recognition

Improving Visual Speech Enhancement Network by Learning Audio-visual Affinity with Multi-head Attention

2022-06-30 · Xinmeng Xu, Yang Wang, Jie Jia, Binbin Chen 외

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution ne…

DecoderSpeech Enhancement