paper-with-me

홈 › Papers

Mixture Encoder for Joint Speech Separation and Recognition

2023-06-21 · Simon Berger, Peter Vieting, Christoph Boeddeker, Ralf Schlüter, Reinhold Haeb-Umbach

Multi-speaker automatic speech recognition (ASR) is crucial for many real-world applications, but it requires dedicated modeling techniques. Existing approaches can be divided into modular and end-to-end methods. Modular approaches separate speakers and recognize each of them with a single-speaker ASR system. End-to-end models process overlapped speech directly in a single, powerful neural network. This work proposes a middle-ground approach that leverages explicit speech separation similarly to the modular approach but also incorporates mixture speech information directly into the ASR module in order to mitigate the propagation of errors made by the speech separator. We also explore a way to exchange cross-speaker context information through a layer that combines information of the individual speakers. Our system is optimized through separate and joint training stages and achieves a relative improvement of 7% in word error rate over a purely modular setup on the SMS-WSJ task.

📄 PDF Abstract BibTeX arXiv:2306.12173

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription

2023-09-15 · Peter Vieting, Simon Berger, Thilo von Neumann, Christoph Boeddeker 외

Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recentl…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Neural Blind Source Separation and Diarization for Distant Speech Recognition

2024-06-12 · Yoshiaki Bando, Tomohiko Nakamura, Shinji Watanabe

This paper presents a neural method for distant speech recognition (DSR) that jointly separates and diarizes speech mixtures without supervision by isolated signals. A standard separation method for multi-talker DSR is a…

blind source separationDistant Speech Recognitionspeaker-diarizationSpeaker Diarization+2

Audio-visual End-to-end Multi-channel Speech Separation, Dereverberation and Recognition

2023-07-06 · Guinan Li, Jiajun Deng, Mengzhe Geng, Zengrui Jin 외

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corrupti…

Speech DereverberationSpeech EnhancementSpeech Separation

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

2026-06-01 · Louis Mouchon arxiv

We present Echo, a proof-of-concept audio system built around a single 25 M-parameter ViT encoder. The encoder is pretrained with a JEPA objective and then specialised by stages to carry speaker identity, phonetic conten…

Speaker DiarizationSpeech Recognition

Multi-talker ASR for an unknown number of sources: Joint training of source counting, separation and ASR

2020-06-04 · Thilo von Neumann, Christoph Boeddeker, Lukas Drude, Keisuke Kinoshita 외

Most approaches to multi-talker overlapped speech separation and recognition assume that the number of simultaneously active speakers is given, but in realistic situations, it is typically unknown. To cope with this, we …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Extractionspeech-recognition+2