paper-with-me

홈 › Papers

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

2025-06-15 · Yuta Hirano, Sakriani Sakti

We propose Speaker-Conditioned Serialized Output Training (SC-SOT), an enhanced SOT-based training for E2E multi-talker ASR. We first probe how SOT handles overlapped speech, and we found the decoder performs implicit speaker separation. We hypothesize this implicit separation is often insufficient due to ambiguous acoustic cues in overlapping regions. To address this, SC-SOT explicitly conditions the decoder on speaker information, providing detailed information about "who spoke when". Specifically, we enhance the decoder by incorporating: (1) speaker embeddings, which allow the model to focus on the acoustic characteristics of the target speaker, and (2) speaker activity information, which guides the model to suppress non-target speakers. The speaker embeddings are derived from a jointly trained E2E speaker diarization model, mitigating the need for speaker enrollment. Experimental results demonstrate the effectiveness of our conditioning approach on overlapped speech.

📄 PDF Abstract BibTeX arXiv:2506.12672

Code (0)

등록된 구현이 없습니다.

Tasks

Decoderspeaker-diarizationSpeaker DiarizationSpeaker Separationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers

2020-06-19 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech. Our model is built on serialized…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSpeaker Identification+2

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning

2019-08-13 · Pavel Denisov, Ngoc Thang Vu

This paper presents our latest investigation on end-to-end automatic speech recognition (ASR) for overlapped speech. We propose to train an end-to-end system conditioned on speaker embeddings and further improved by tran…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Serialized Output Training for End-to-End Overlapped Speech Recognition

2020-03-28 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instead of having multiple output layers as wi…

Decoderspeech-recognitionSpeech Recognition

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

2025-09-18 · Miseul Kim, Soo Jin Park, Kyungguen Byun, Hyeon-Kyeong Shin 외 arxiv

Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different indi…

Speaker Diarization

MIRNet: Learning multiple identities representations in overlapped speech

2020-08-04 · Hyewon Han, Soo-Whan Chung, Hong-Goo Kang

Many approaches can derive information about a single speaker's identity from the speech by learning to recognize consistent characteristics of acoustic parameters. However, it is challenging to determine identity inform…

Rgb-T TrackingSpeaker VerificationSpeech Separation