paper-with-me

Papers

Low-Latency Speaker-Independent Continuous Speech Separation

2019-04-13 · Takuya Yoshioka, Zhuo Chen, Changliang Liu, Xiong Xiao, Hakan Erdogan, Dimitrios Dimitriadis

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of which contains no overlapping speech segment. A separated, or cleaned, version of each utterance is generated from one of SI-CSS's output channels nondeterministically without being split up and distributed to multiple channels. A typical application scenario is transcribing multi-party conversations, such as meetings, recorded with microphone arrays. The output signals can be simply sent to a speech recognition engine because they do not include speech overlaps. The previous SI-CSS method uses a neural network trained with permutation invariant training and a data-driven beamformer and thus requires much processing latency. This paper proposes a low-latency SI-CSS method whose performance is comparable to that of the previous method in a microphone array-based meeting transcription task.This is achieved (1) by using a new speech separation network architecture combined with a double buffering scheme and (2) by performing enhancement with a set of fixed beamformers followed by a neural post-filter.

📄 PDF Abstract BibTeX arXiv:1904.06478

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation

2018-09-20 · Yi Luo, Nima Mesgarani

Single-channel, speaker-independent speech separation methods have recently seen great progress. However, the accuracy, latency, and computational cost of such methods remain insufficient. The majority of the previous me…

Multi-task Audio Source SeperationMusic Source SeparationSpeaker SeparationSpeech Enhancement+1

SkiM: Skipping Memory LSTM for Low-Latency Real-Time Continuous Speech Separation

2022-01-26 · Chenda Li, Lei Yang, Weiqin Wang, Yanmin Qian

Continuous speech separation for meeting pre-processing has recently become a focused research topic. Compared to the data in utterance-level speech separation, the meeting-style audio stream lasts longer, has an uncerta…

Speech Separation

Deep neural network Based Low-latency Speech Separation with Asymmetric analysis-Synthesis Window Pair

2021-06-22 · Shanshan Wang, Gaurav Naithani, Archontis Politis, Tuomas Virtanen

Time-frequency masking or spectrum prediction computed via short symmetric windows are commonly used in low-latency deep neural network (DNN) based source separation. In this paper, we propose the usage of an asymmetric …

ClusteringDeep ClusteringSpeech EnhancementSpeech Separation

Multi-channel Conversational Speaker Separation via Neural Diarization

2023-11-15 · Hassan Taherian, DeLiang Wang

When dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or mee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognition+1

Conversational Speech Separation: an Evaluation Study for Streaming Applications

2022-05-31 · Giovanni Morrone, Samuele Cornell, Enrico Zovato, Alessio Brutti 외

Continuous speech separation (CSS) is a recently proposed framework which aims at separating each speaker from an input mixture signal in a streaming fashion. Hereafter we perform an evaluation study on practical design …

Speech Separation