paper-with-me

Papers

Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech

2019-05-09 · Tobias Menne, Ilya Sklyar, Ralf Schlüter, Hermann Ney

Significant performance degradation of automatic speech recognition (ASR) systems is observed when the audio signal contains cross-talk. One of the recently proposed approaches to solve the problem of multi-speaker ASR is the deep clustering (DPCL) approach. Combining DPCL with a state-of-the-art hybrid acoustic model, we obtain a word error rate (WER) of 16.5 % on the commonly used wsj0-2mix dataset, which is the best performance reported thus far to the best of our knowledge. The wsj0-2mix dataset contains simulated cross-talk where the speech of multiple speakers overlaps for almost the entire utterance. In a more realistic ASR scenario the audio signal contains significant portions of single-speaker speech and only part of the signal contains speech of multiple competing speakers. This paper investigates obstacles of applying DPCL as a preprocessing method for ASR in such a scenario of sparsely overlapping speech. To this end we present a data simulation approach, closely related to the wsj0-2mix dataset, generating sparsely overlapping speech datasets of arbitrary overlap ratio. The analysis of applying DPCL to sparsely overlapping speech is an important interim step between the fully overlapping datasets like wsj0-2mix and more realistic ASR datasets, such as CHiME-5 or AMI.

📄 PDF Abstract BibTeX arXiv:1905.03500

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringDeep Clusteringspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Discrete Speech Unit Extraction via Independent Component Analysis

2025-01-11 · Tomohiko Nakamura, Kwanghee Choi, Keigo Hojo, Yoshiaki Bando 외

Self-supervised speech models (S3Ms) have become a common tool for the speech processing community, leveraging representations for downstream tasks. Clustering S3M representations yields discrete speech units (DSUs), whi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeech-recognition+1

AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data

2023-09-25 · Jianwei Yu, Hangting Chen, Yanyao Bian, Xiang Li 외

Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech t…

Automatic Speech RecognitionSpeech EnhancementSpeech Extractionspeech-recognition+1

KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition

2020-09-07 · Soohwan Kim, Seyoung Bae, Cheolhwang Won

We present KoSpeech, an open-source software, which is modular and extensible end-to-end Korean automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch. Several automatic speech recognition …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Unsupervised recognition and clustering of speech overlaps in spoken conversations

2014-09-11 · Workshop on Speech, Language and Audio in Multimedia (SLAM 2014) 2014 9 · Shammur Absar Chowdhury, Giuseppe Riccardi, Firoj Alam

We are interested in understanding speech overlaps and their function in human conversations. Previous studies on speech overlaps have relied on supervised methods, small corpora and controlled conversations. The charact…

ClusteringSpeech Interruption Detection

When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper

2026-03-05 · Akif Islam, Raufun Nahar, Md. Ekramul Hamid arxiv

Recent advances in automatic speech recognition (ASR) and speech enhancement have led to a widespread assumption that improving perceptual audio quality should directly benefit recognition accuracy. In this work, we rigo…

Speech RecognitionSpeech Enhancement