paper-with-me

홈 › Papers

Simulating realistic speech overlaps improves multi-talker ASR

2022-10-27 · Muqiao Yang, Naoyuki Kanda, Xiaofei Wang, Jian Wu, Sunit Sivasankaran, Zhuo Chen, Jinyu Li, Takuya Yoshioka

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the difficulty in acquiring real conversation data with high-quality human transcriptions, a na\"ive simulation of multi-talker speech by randomly mixing multiple utterances was conventionally used for model training. In this work, we propose an improved technique to simulate multi-talker overlapping speech with realistic speech overlaps, where an arbitrary pattern of speech overlaps is represented by a sequence of discrete tokens. With this representation, speech overlapping patterns can be learned from real conversations based on a statistical language model, such as N-gram, which can be then used to generate multi-talker speech for training. In our experiments, multi-talker ASR models trained with the proposed method show consistent improvement on the word error rates across multiple datasets.

📄 PDF Abstract BibTeX arXiv:2210.15715

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech

2019-05-09 · Tobias Menne, Ilya Sklyar, Ralf Schlüter, Hermann Ney

Significant performance degradation of automatic speech recognition (ASR) systems is observed when the audio signal contains cross-talk. One of the recently proposed approaches to solve the problem of multi-speaker ASR i…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringDeep Clustering+2

Unsupervised recognition and clustering of speech overlaps in spoken conversations

2014-09-11 · Workshop on Speech, Language and Audio in Multimedia (SLAM 2014) 2014 9 · Shammur Absar Chowdhury, Giuseppe Riccardi, Firoj Alam

We are interested in understanding speech overlaps and their function in human conversations. Previous studies on speech overlaps have relied on supervised methods, small corpora and controlled conversations. The charact…

ClusteringSpeech Interruption Detection

Recognizing long-form speech using streaming end-to-end models

2019-10-24 · Arun Narayanan, Rohit Prabhavalkar, Chung-Cheng Chiu, David Rybach 외

All-neural end-to-end (E2E) automatic speech recognition (ASR) systems that use a single neural network to transduce audio to word sequences have been shown to achieve state-of-the-art results on several tasks. In this w…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityForm+2

FRA-RIR: Fast Random Approximation of the Image-source Method

2022-08-08 · Yi Luo, Jianwei Yu

The training of modern speech processing systems often requires a large amount of simulated room impulse response (RIR) data in order to allow the systems to generalize well in real-world, reverberant environments. Howev…

DenoisingGPURoom Impulse Response (RIR)Speech Denoising

Simulating dysarthric speech for training data augmentation in clinical speech applications

2018-04-27

Training machine learning algorithms for speech applications requires large, labeled training data sets. This is problematic for clinical applications where obtaining such data is prohibitively expensive because of priva…

Data Augmentation