paper-with-me

Papers

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

2025-05-25 · Helin Wang, Jiarui Hai, Dongchao Yang, Chen Chen, Kai Li, Junyi Peng, Thomas Thebaud, Laureano Moro Velazquez, Jesus Villalba, Najim Dehak

Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent advancements in TSE have primarily employed discriminative models that offer high perceptual quality, these models often introduce unwanted artifacts, reduce naturalness, and are sensitive to discrepancies between training and testing environments. On the other hand, generative models for TSE lag in perceptual quality and intelligibility. To address these challenges, we present SoloSpeech, a novel cascaded generative pipeline that integrates compression, extraction, reconstruction, and correction processes. SoloSpeech features a speaker-embedding-free target extractor that utilizes conditional information from the cue audio's latent space, aligning it with the mixture audio's latent space to prevent mismatches. Evaluated on the widely-used Libri2Mix dataset, SoloSpeech achieves the new state-of-the-art intelligibility and quality in target speech extraction and speech separation tasks while demonstrating exceptional generalization on out-of-domain data and real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2505.19314

Code (1)

wanghelin1997/solospeech 공식 구현 pytorch

Tasks

Speech ExtractionSpeech Separation

Similar Papers 제목 키워드 기반

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

2025-01-24 · Hao Ma, Rujin Chen, Xiao-Lei Zhang, Ju Liu 외

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spect…

Speech Extraction

Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise

2022-03-20 · Tuomo Raitio, Petko Petkov, Jiangchuan Li, Muhammed Shifas 외

We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise. The method consists of first measuring the spectral …

text-to-speechText to Speech

Minimum Processing Near-end Listening Enhancement

2022-10-31 · Andreas Jonas Fuglsig, Jesper Jensen, Zheng-Hua Tan, Lars Søndergaard Bertelsen 외

The intelligibility and quality of speech from a mobile phone or public announcement system are often affected by background noise in the listening environment. By pre-processing the speech signal it is possible to impro…

Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility

2024-09-14 · Xiaoyu Liu, Xu Li, Joan Serrà, Santiago Pascual

Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its …

Knowledge DistillationLanguage ModelingLanguage Modelling

ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams

2024-10-23 · Srija Anand, Praveen Srinivasa Varadhan, Mehak Singal, Mitesh M. Khapra

Recent advancements in Text-to-Speech (TTS) technology have led to natural-sounding speech for English, primarily due to the availability of large-scale, high-quality web data. However, many other languages lack access t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingKnowledge Distillation+5