paper-with-me

Papers

Audio-visual End-to-end Multi-channel Speech Separation, Dereverberation and Recognition

2023-07-06 · Guinan Li, Jiajun Deng, Mengzhe Geng, Zengrui Jin, Tianzi Wang, Shujie Hu, Mingyu Cui, Helen Meng, Xunying Liu

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corruption, an audio-visual multi-channel speech separation, dereverberation and recognition approach featuring a full incorporation of visual information into all system components is proposed in this paper. The efficacy of the video input is consistently demonstrated in mask-based MVDR speech separation, DNN-WPE or spectral mapping (SpecM) based speech dereverberation front-end and Conformer ASR back-end. Audio-visual integrated front-end architectures performing speech separation and dereverberation in a pipelined or joint fashion via mask-based WPD are investigated. The error cost mismatch between the speech enhancement front-end and ASR back-end components is minimized by end-to-end jointly fine-tuning using either the ASR cost function alone, or its interpolation with the speech enhancement loss. Experiments were conducted on the mixture overlapped and reverberant speech data constructed using simulation or replay of the Oxford LRS2 dataset. The proposed audio-visual multi-channel speech separation, dereverberation and recognition systems consistently outperformed the comparable audio-only baseline by 9.1% and 6.2% absolute (41.7% and 36.0% relative) word error rate (WER) reductions. Consistent speech enhancement improvements were also obtained on PESQ, STOI and SRMR scores.

📄 PDF Abstract BibTeX arXiv:2307.02909

Code (0)

등록된 구현이 없습니다.

Tasks

Speech DereverberationSpeech EnhancementSpeech Separation

Similar Papers 제목 키워드 기반

Audio-visual multi-channel speech separation, dereverberation and recognition

2022-04-05 · Guinan Li, Jianwei Yu, Jiajun Deng, Xunying Liu 외

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverbera…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+2

Audio-visual Multi-channel Recognition of Overlapped Speech

2020-05-18 · Jianwei Yu, Bo Wu, Rongzhi Gu, Shi-Xiong Zhang 외

Automatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-art ASR systems. Motivated by the invari…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)LipreadingSentence+3

Audio-visual Multi-channel Integration and Recognition of Overlapped Speech

2020-11-16 · Jianwei Yu, Shi-Xiong Zhang, Bo Wu, Shansong Liu 외

Automatic speech recognition (ASR) technologies have been significantly advanced in the past few decades. However, recognition of overlapped speech remains a highly challenging task to date. To this end, multi-channel mi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

FaceFilter: Audio-visual speech separation using still images

2020-05-14 · Soo-Whan Chung, Soyeon Choe, Joon Son Chung, Hong-Goo Kang

The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-…

Speech Separation

Deep Variational Generative Models for Audio-visual Speech Separation

2020-08-17 · Viet-Nhat Nguyen, Mostafa Sadeghi, Elisa Ricci, Xavier Alameda-Pineda

In this paper, we are interested in audio-visual speech separation given a single-channel audio recording as well as visual information (lips movements) associated with each speaker. We propose an unsupervised technique …

Speech Separation