paper-with-me

홈 › Papers

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio

2025-05-16 · Xinlu He, Jacob Whitehill

Monaural multi-speaker automatic speech recognition (ASR) remains challenging due to data scarcity and the intrinsic difficulty of recognizing and attributing words to individual speakers, particularly in overlapping speech. Recent advances have driven the shift from cascade systems to end-to-end (E2E) architectures, which reduce error propagation and better exploit the synergy between speech content and speaker identity. Despite rapid progress in E2E multi-speaker ASR, the field lacks a comprehensive review of recent developments. This survey provides a systematic taxonomy of E2E neural approaches for multi-speaker ASR, highlighting recent advances and comparative analysis. Specifically, we analyze: (1) architectural paradigms (SIMO vs.~SISO) for pre-segmented audio, analyzing their distinct characteristics and trade-offs; (2) recent architectural and algorithmic improvements based on these two paradigms; (3) extensions to long-form speech, including segmentation strategy and speaker-consistent hypothesis stitching. Further, we (4) evaluate and compare methods across standard benchmarks. We conclude with a discussion of open challenges and future research directions towards building robust and scalable multi-speaker ASR.

📄 PDF Abstract BibTeX arXiv:2505.10975

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Deep Neural Networks for Automatic Speech Processing: A Survey from Large Corpora to Limited Data

2020-03-09 · Vincent Roger, Jérôme Farinas, Julien Pinquier

Most state-of-the-art speech systems are using Deep Neural Networks (DNNs). Those systems require a large amount of data to be learned. Hence, learning state-of-the-art frameworks on under-resourced speech languages/prob…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeaker Identification+2

Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers

2020-06-19 · Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng 외

We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech. Our model is built on serialized…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSpeaker Identification+2

Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends

2020-01-02 · Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak 외

Research on speech processing has traditionally considered the task of designing hand-engineered acoustic features (feature engineering) as a separate distinct problem from the task of designing efficient machine learnin…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionFeature Engineering+6

End-to-End Monaural Multi-speaker ASR System without Pretraining

2018-11-05 · Xuankai Chang, Yanmin Qian, Kai Yu, Shinji Watanabe

Recently, end-to-end models have become a popular approach as an alternative to traditional hybrid models in automatic speech recognition (ASR). The multi-speaker speech separation and recognition task is a central task …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning

2019-08-13 · Pavel Denisov, Ngoc Thang Vu

This paper presents our latest investigation on end-to-end automatic speech recognition (ASR) for overlapped speech. We propose to train an end-to-end system conditioned on speaker embeddings and further improved by tran…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1