paper-with-me

홈 › Papers

Audio-visual speech separation based on joint feature representation with cross-modal attention

2022-03-05 · Junwen Xiong, Peng Zhang, Lei Xie, Wei Huang, Yufei zha, Yanning Zhang

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on feature learning of each single modality, which is far from sufficient consideration of inter-relationships between modalites. Inspired by learning joint feature representations from audio and visual streams with attention mechanism, in this study, a novel cross-modal fusion strategy is proposed to benefit the whole framework with semantic correlations between different modalities. To further improve audio-visual speech separation, the dense optical flow of lip motion is incorporated to strengthen the robustness of visual representation. The evaluation of the proposed work is performed on two public audio-visual speech separation benchmark datasets. The overall improvement of the performance has demonstrated that the additional motion network effectively enhances the visual representation of the combined lip images and audio signal, as well as outperforming the baseline in terms of all metrics with the proposed cross-modal fusion.

📄 PDF Abstract BibTeX arXiv:2203.02655

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationSpeech Separation

Similar Papers 제목 키워드 기반

Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

2018-04-10 · Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel 외

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and do…

Speech Separation

VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency

2021-01-08 · CVPR 2021 1 · Ruohan Gao, Kristen Grauman

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas exis…

Speech Separation

Audio-visual End-to-end Multi-channel Speech Separation, Dereverberation and Recognition

2023-07-06 · Guinan Li, Jiajun Deng, Mengzhe Geng, Zengrui Jin 외

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corrupti…

Speech DereverberationSpeech EnhancementSpeech Separation

RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation

2023-09-29 · Samuel Pegg, Kai Li, Xiaolin Hu

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing stat…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Separation+1

Audio-visual Speech Separation with Adversarially Disentangled Visual Representation

2020-11-29 · Peng Zhang, Jiaming Xu, Jing Shi, Yunzhe Hao 외

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they build on a strategy to handle the predefin…

Speech Separation