paper-with-me

홈 › Papers

A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio

2021-07-06 · Naoyuki Kanda, Xiong Xiao, Jian Wu, Tianyan Zhou, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, Takuya Yoshioka

Speaker-attributed automatic speech recognition (SA-ASR) is a task to recognize "who spoke what" from multi-talker recordings. An SA-ASR system usually consists of multiple modules such as speech separation, speaker diarization and ASR. On the other hand, considering the joint optimization, an end-to-end (E2E) SA-ASR model has recently been proposed with promising results on simulation data. In this paper, we present our recent study on the comparison of such modular and joint approaches towards SA-ASR on real monaural recordings. We develop state-of-the-art SA-ASR systems for both modular and joint approaches by leveraging large-scale training data, including 75 thousand hours of ASR training data and the VoxCeleb corpus for speaker representation learning. We also propose a new pipeline that performs the E2E SA-ASR model after speaker clustering. Our evaluation on the AMI meeting corpus reveals that after fine-tuning with a small real data, the joint system performs 8.9--29.9% better in accuracy compared to the best modular system while the modular system performs better before such fine-tuning. We also conduct various error analyses to show the remaining issues for the monaural SA-ASR.

📄 PDF Abstract BibTeX arXiv:2107.02852

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)FormRepresentation Learningspeaker-diarizationSpeaker Diarizationspeech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party Meetings

2022-03-31 · Fan Yu, Zhihao Du, Shiliang Zhang, Yuxiao Lin 외

In this paper, we conduct a comparative study on speaker-attributed automatic speech recognition (SA-ASR) in the multi-party meeting scenario, a topic with increasing attention in meeting rich transcription. Specifically…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognition+1

Mixture Encoder for Joint Speech Separation and Recognition

2023-06-21 · Simon Berger, Peter Vieting, Christoph Boeddeker, Ralf Schlüter 외

Multi-speaker automatic speech recognition (ASR) is crucial for many real-world applications, but it requires dedicated modeling techniques. Existing approaches can be divided into modular and end-to-end methods. Modular…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

A Comparative Study on Multichannel Speaker-Attributed Automatic Speech Recognition in Multi-party Meetings

2022-11-01 · Mohan Shi, Jie Zhang, Zhihao Du, Fan Yu 외

Speaker-attributed automatic speech recognition (SA-ASR) in multi-party meeting scenarios is one of the most valuable and challenging ASR task. It was shown that single-channel frame-level diarization with serialized out…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization

2024-11-04 · Petr Pálka, Federico Landini, Dominik Klement, Mireia Diez 외

In spite of the popularity of end-to-end diarization systems nowadays, modular systems comprised of voice activity detection (VAD), speaker embedding extraction plus clustering, and overlapped speech detection (OSD) plus…

Action DetectionActivity DetectionClustering

A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR

2024-09-09 · Giovanni Morrone, Enrico Zovato, Fabio Brugnara, Enrico Sartori 외

We present a modular toolkit to perform joint speaker diarization and speaker identification. The toolkit can leverage on multiple models and algorithms which are defined in a configuration file. Such flexibility allows …

Automatic Speech Recognitionspeaker-diarizationSpeaker DiarizationSpeaker Identification+2