paper-with-me

Papers

Joint ASR and Speaker Role Tagging with Serialized Output Training

2025-06-12 · Anfeng Xu, Tiantian Feng, Shrikanth Narayanan

Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function that is critical for conversational AI. In this work, we investigate the use of serialized output training (SOT) for joint ASR and speaker role tagging. By augmenting Whisper with role-specific tokens and fine-tuning it with SOT, we enable the model to generate role-aware transcriptions in a single decoding pass. We compare the SOT approach against a self-supervised previous baseline method on two real-world conversational datasets. Our findings show that this approach achieves more than 10% reduction in multi-talker WER, demonstrating its feasibility as a unified model for speaker-role aware speech transcription.

📄 PDF Abstract BibTeX arXiv:2506.10349

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음
AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

2025-10-04 · Martin Kocour, Martin Karafiat, Alexander Polok, Dominik Klement 외 arxiv

We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned W…

Speech Recognition

BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR

2023-05-23 · Yuhao Liang, Fan Yu, Yangze Li, Pengcheng Guo 외

The recently proposed serialized output training (SOT) simplifies multi-talker automatic speech recognition (ASR) by generating speaker transcriptions separated by a special token. However, frequent speaker changes can m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Change DetectionDecoder+2

Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers

2020-06-19 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech. Our model is built on serialized…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSpeaker Identification+2

One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition

2023-10-02 · Samuele Cornell, Jee-weon Jung, Shinji Watanabe, Stefano Squartini

This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoder+4

Serialized Output Training for End-to-End Overlapped Speech Recognition

2020-03-28 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instead of having multiple output layers as wi…

Decoderspeech-recognitionSpeech Recognition