paper-with-me

Papers

Conversational Speech Separation: an Evaluation Study for Streaming Applications

2022-05-31 · Giovanni Morrone, Samuele Cornell, Enrico Zovato, Alessio Brutti, Stefano Squartini

Continuous speech separation (CSS) is a recently proposed framework which aims at separating each speaker from an input mixture signal in a streaming fashion. Hereafter we perform an evaluation study on practical design considerations for a CSS system, addressing important aspects which have been neglected in recent works. In particular, we focus on the trade-off between separation performance, computational requirements and output latency showing how an offline separation algorithm can be used to perform CSS with a desired latency. We carry out an extensive analysis on the choice of CSS processing window size and hop size on sparsely overlapped data. We find out that the best trade-off between computational burden and performance is obtained for a window of 5 s.

📄 PDF Abstract BibTeX arXiv:2205.15700

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

2022-09-12 · Naoyuki Kanda, Jian Wu, Xiaofei Wang, Zhuo Chen 외

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, cap…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Separation Guided Speaker Diarization in Realistic Mismatched Conditions

2021-07-06 · Shu-Tong Niu, Jun Du, Lei Sun, Chin-Hui Lee

We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clustering-based speaker diarization (CSD) appro…

Clusteringspeaker-diarizationSpeaker DiarizationSpeech Separation

End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations

2023-03-21 · Giovanni Morrone, Samuele Cornell, Luca Serafini, Enrico Zovato 외

Recent works show that speech separation guided diarization (SSGD) is an increasingly promising direction, mainly thanks to the recent progress in speech separation. It performs diarization by first separating the speake…

Action DetectionActivity DetectionAutomatic Speech Recognitionspeech-recognition+2

SpeakStream: Streaming Text-to-Speech with Interleaved Data

2025-05-25 · Richard He Bai, Zijin Gu, Tatiana Likhomanenko, Navdeep Jaitly

The latency bottleneck of traditional text-to-speech (TTS) systems fundamentally hinders the potential of streaming large language models (LLMs) in conversational AI. These TTS systems, typically trained and inferenced o…

Decodertext-to-speechText to Speech

VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition

2020-09-09 · Quan Wang, Ignacio Lopez Moreno, Mert Saglam, Kevin Wilson 외

We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speech recognition system. Delivering such a …

CPUspeech-recognitionSpeech Recognition