paper-with-me

Papers

Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis

2020-11-03 · Desh Raj, Pavel Denisov, Zhuo Chen, Hakan Erdogan, Zili Huang, Maokui He, Shinji Watanabe, Jun Du, Takuya Yoshioka, Yi Luo, Naoyuki Kanda, Jinyu Li, Scott Wisdom, John R. Hershey

Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in systems dealing with speech separation, speaker diarization, and automatic speech recognition (ASR) in the last decade, it has become possible to build pipelines that achieve reasonable error rates on this task. In this paper, we propose an end-to-end modular system for the LibriCSS meeting data, which combines independently trained separation, diarization, and recognition components, in that order. We study the effect of different state-of-the-art methods at each stage of the pipeline, and report results using task-specific metrics like SDR and DER, as well as downstream WER. Experiments indicate that the problem of overlapping speech for diarization and ASR can be effectively mitigated with the presence of a well-trained separation module. Our best system achieves a speaker-attributed WER of 12.7%, which is close to that of a non-overlapping ASR.

📄 PDF Abstract BibTeX arXiv:2011.02014

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations

2023-03-21 · Giovanni Morrone, Samuele Cornell, Luca Serafini, Enrico Zovato 외

Recent works show that speech separation guided diarization (SSGD) is an increasingly promising direction, mainly thanks to the recent progress in speech separation. It performs diarization by first separating the speake…

Action DetectionActivity DetectionAutomatic Speech Recognitionspeech-recognition+2

DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions

2024-11-11 · Shu-Tong Niu, Jun Du, Ruo-Yu Wang, Gao-Bin Yang 외

We propose a single-channel Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework for back-end automatic speech recognition (ASR), combining neural speaker diarization (NSD) and speech separation (SS). Fir…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+4

Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization

2023-09-28 · Thilo von Neumann, Christoph Boeddeker, Tobias Cord-Landwehr, Marc Delcroix 외

We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuous Speech Separation (CSS) system with a…

SentenceSpeech Separation

Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models

2024-10-28 · Tobias Cord-Landwehr, Christoph Boeddeker, Reinhold Haeb-Umbach

We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Mo…

Speech Enhancement

Multi-channel Conversational Speaker Separation via Neural Diarization

2023-11-15 · Hassan Taherian, DeLiang Wang

When dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or mee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognition+1