Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
This paper presents a novel framework for multi-talker automatic speech recognition without the need for auxiliary information. Serialized Output Training (SOT), a widely used approach, suffers from recognition errors due to speaker assignment failures. Although incorporating auxiliary information, such as token-level timestamps, can improve recognition accuracy, extracting such information from natural conversational speech remains challenging. To address this limitation, we propose Speaker-Distinguishable CTC (SD-CTC), an extension of CTC that jointly assigns a token and its corresponding speaker label to each frame. We further integrate SD-CTC into the SOT framework, enabling the SOT model to learn speaker distinction using only overlapping speech and transcriptions. Experimental comparisons show that multi-task learning with SD-CTC and SOT reduces the error rate of the SOT model by 26% and achieves performance comparable to state-of-the-art methods relying on auxiliary information.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionMulti-Task Learningspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Target Speaker Verification with Selective Auditory Attention for Single and Multi-talker Speech
Speaker verification has been studied mostly under the single-talker condition. It is adversely affected in the presence of interference speakers. Inspired by the study on target speaker extraction, e.g., SpEx, we propos…
Multi-Task LearningSpeaker VerificationTarget Speaker ExtractionSpeaker Mask Transformer for Multi-talker Overlapped Speech Recognition
Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper, to better address these tasks, we first…
speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech RecognitionUSEV: Universal Speaker Extraction with Visual Cue
A speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. H…
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned W…
Speech RecognitionUnified Autoregressive Modeling for Joint End-to-End Multi-Talker Overlapped Speech Recognition and Speaker Attribute Estimation
In this paper, we present a novel modeling method for single-channel multi-talker overlapped automatic speech recognition (ASR) systems. Fully neural network based end-to-end models have dramatically improved the perform…
Age EstimationAttributeAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+2