paper-with-me

Papers

Exploring the Integration of Speech Separation and Recognition with Self-Supervised Learning Representation

2023-07-23 · Yoshiki Masuyama, Xuankai Chang, Wangyou Zhang, Samuele Cornell, Zhong-Qiu Wang, Nobutaka Ono, Yanmin Qian, Shinji Watanabe

Neural speech separation has made remarkable progress and its integration with automatic speech recognition (ASR) is an important direction towards realizing multi-speaker ASR. This work provides an insightful investigation of speech separation in reverberant and noisy-reverberant scenarios as an ASR front-end. In detail, we explore multi-channel separation methods, mask-based beamforming and complex spectral mapping, as well as the best features to use in the ASR back-end model. We employ the recent self-supervised learning representation (SSLR) as a feature and improve the recognition performance from the case with filterbank features. To further improve multi-speaker recognition performance, we present a carefully designed training strategy for integrating speech separation and recognition with SSLR. The proposed integration using TF-GridNet-based complex spectral mapping and WavLM-based SSLR achieves a 2.5% word error rate in reverberant WHAMR! test set, significantly outperforming an existing mask-based MVDR beamforming and filterbank integration (28.9%).

📄 PDF Abstract BibTeX arXiv:2307.12231

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised LearningSpeaker Recognitionspeech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

ESPnet-se: end-to-end speech enhancement and separation toolkit designed for asr integration

2020-11-07

We present ESPnet-SE, which is designed for the quick development of speech enhancement and speech separation systems in a single framework, along with the optional downstream speech recognition module. ESPnet-SE is a ne…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingSpeech Enhancement+3

Audio-visual multi-channel speech separation, dereverberation and recognition

2022-04-05 · Guinan Li, Jianwei Yu, Jiajun Deng, Xunying Liu 외

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverbera…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+2

Audio-visual Multi-channel Integration and Recognition of Overlapped Speech

2020-11-16 · Jianwei Yu, Shi-Xiong Zhang, Bo Wu, Shansong Liu 외

Automatic speech recognition (ASR) technologies have been significantly advanced in the past few decades. However, recognition of overlapped speech remains a highly challenging task to date. To this end, multi-channel mi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis

2020-11-03 · Desh Raj, Pavel Denisov, Zhuo Chen, Hakan Erdogan 외

Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in systems dealing with speech separation, spea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

Exploring Self-Attention Mechanisms for Speech Separation

2022-02-06 · Cem Subakan, Mirco Ravanelli, Samuele Cornell, Francois Grondin 외

Transformers have enabled impressive improvements in deep learning. They often outperform recurrent and convolutional models in many tasks while taking advantage of parallel processing. Recently, we proposed the SepForme…

DenoisingSpeech EnhancementSpeech Separation