paper-with-me

Papers

Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription

2023-09-15 · Peter Vieting, Simon Berger, Thilo von Neumann, Christoph Boeddeker, Ralf Schlüter, Reinhold Haeb-Umbach

Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement.

📄 PDF Abstract BibTeX arXiv:2309.08454

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

2023-10-30 · Zexu Pan, Gordon Wichern, Yoshiki Masuyama, Francois G. Germain 외

Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. Building upon the achievements of the sta…

Speaker SeparationSpeech EnhancementSpeech Extraction

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

2025-02-03 · Weiwei Lin, Chenghan He

We propose a novel autoregressive modeling approach for speech synthesis, combining a variational autoencoder (VAE) with a multi-modal latent space and an autoregressive model that uses Gaussian Mixture Models (GMM) as t…

QuantizationSpeech Synthesis

Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization

2023-09-28 · Thilo von Neumann, Christoph Boeddeker, Tobias Cord-Landwehr, Marc Delcroix 외

We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuous Speech Separation (CSS) system with a…

SentenceSpeech Separation

Multi-Channel Target Speaker Extraction with Refinement: The WavLab Submission to the Second Clarity Enhancement Challenge

2023-02-15 · Samuele Cornell, Zhong-Qiu Wang, Yoshiki Masuyama, Shinji Watanabe 외

This paper describes our submission to the Second Clarity Enhancement Challenge (CEC2), which consists of target speech enhancement for hearing-aid (HA) devices in noisy-reverberant environments with multiple interferers…

Speaker SeparationSpeech EnhancementTarget Speaker Extraction

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

2025-05-27 · Zexu Pan, Shengkui Zhao, Tingting Wang, Kun Zhou 외

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co…