paper-with-me

Papers

BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR

2023-05-23 · Yuhao Liang, Fan Yu, Yangze Li, Pengcheng Guo, Shiliang Zhang, Qian Chen, Lei Xie

The recently proposed serialized output training (SOT) simplifies multi-talker automatic speech recognition (ASR) by generating speaker transcriptions separated by a special token. However, frequent speaker changes can make speaker change prediction difficult. To address this, we propose boundary-aware serialized output training (BA-SOT), which explicitly incorporates boundary knowledge into the decoder via a speaker change detection task and boundary constraint loss. We also introduce a two-stage connectionist temporal classification (CTC) strategy that incorporates token-level SOT CTC to restore temporal context information. Besides typical character error rate (CER), we introduce utterance-dependent character error rate (UD-CER) to further measure the precision of speaker change prediction. Compared to original SOT, BA-SOT reduces CER/UD-CER by 5.1%/14.0%, and leveraging a pre-trained ASR model for BA-SOT model initialization further reduces CER/UD-CER by 8.4%/19.9%.

📄 PDF Abstract BibTeX arXiv:2305.13716

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Change DetectionDecoderspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Joint ASR and Speaker Role Tagging with Serialized Output Training

2025-06-12 · Anfeng Xu, Tiantian Feng, Shrikanth Narayanan

Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition

2025-09-01 · Hao Shi, Yusuke Fujita, Tomoya Mizumoto, Lianbo Liu 외 arxiv

Prompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic speech recognition (ASR) systems either omit…

Information ExtractionSpeech Recognition

Serialized Output Training for End-to-End Overlapped Speech Recognition

2020-03-28 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instead of having multiple output layers as wi…

Decoderspeech-recognitionSpeech Recognition

Streaming Multi-Talker ASR with Token-Level Serialized Output Training

2022-02-02 · Naoyuki Kanda, Jian Wu, Yu Wu, Xiong Xiao 외

This paper proposes a token-level serialized output training (t-SOT), a novel framework for streaming multi-talker automatic speech recognition (ASR). Unlike existing streaming multi-talker ASR models using multiple outp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Serialized Output Training by Learned Dominance

2024-07-04 · Ying Shi, Lantian Li, Shi Yin, Dong Wang 외

Serialized Output Training (SOT) has showcased state-of-the-art performance in multi-talker speech recognition by sequentially decoding the speech of individual speakers. To address the challenging label-permutation issu…

Decoderspeech-recognitionSpeech Recognition