paper-with-me

홈 › Papers

Stepwise-Refining Speech Separation Network via Fine-Grained Encoding in High-order Latent Domain

2021-10-10 · Zengwei Yao, Wenjie Pei, Fanglin Chen, Guangming Lu, David Zhang

The crux of single-channel speech separation is how to encode the mixture of signals into such a latent embedding space that the signals from different speakers can be precisely separated. Existing methods for speech separation either transform the speech signals into frequency domain to perform separation or seek to learn a separable embedding space by constructing a latent domain based on convolutional filters. While the latter type of methods learning an embedding space achieves substantial improvement for speech separation, we argue that the embedding space defined by only one latent domain does not suffice to provide a thoroughly separable encoding space for speech separation. In this paper, we propose the Stepwise-Refining Speech Separation Network (SRSSN), which follows a coarse-to-fine separation framework. It first learns a 1-order latent domain to define an encoding space and thereby performs a rough separation in the coarse phase. Then the proposed SRSSN learns a new latent domain along each basis function of the existing latent domain to obtain a high-order latent domain in the refining phase, which enables our model to perform a refining separation to achieve a more precise speech separation. We demonstrate the effectiveness of our SRSSN by conducting extensive experiments, including speech separation in a clean (noise-free) setting on WSJ0-2/3mix datasets as well as in noisy/reverberant settings on WHAM!/WHAMR! datasets. Furthermore, we also perform experiments of speech recognition on separated speech signals by our model to evaluate the performance of speech separation indirectly.

📄 PDF Abstract BibTeX arXiv:2110.04791

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization

2025-09-29 · Jiacheng Shi, Hongfei Du, Yangfan He, Y. Alicia Hong 외 arxiv

Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotio…

Text to Speech

EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers

2022-03-31 · Soumi Maiti, Yushi Ueda, Shinji Watanabe, Chunlei Zhang 외

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neura…

Decoderspeaker-diarizationSpeaker DiarizationSpeech Separation

Prompt Chaining or Stepwise Prompt? Refinement in Text Summarization

2024-06-01 · Shichao Sun, Ruifeng Yuan, Ziqiang Cao, Wenjie Li 외

Large language models (LLMs) have demonstrated the capacity to improve summary quality by mirroring a human-like iterative process of critique and refinement starting from the initial draft. Two strategies are designed t…

Text Summarization

Deep Neural Mel-Subband Beamformer for In-car Speech Separation

2022-11-22 · Vinay Kothapally, Yong Xu, Meng Yu, Shi-Xiong Zhang 외

While current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies independently which results in higher computa…

Speech Separation

TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network

2025-05-24 · Xiaobin Rong, DaHan Wang, Qinwen Hu, Yushi Wang 외

Universal speech enhancement aims to handle input speech with different distortions and input formats. To tackle this challenge, we present TS-URGENet, a Three-Stage Universal, Robust, and Generalizable speech Enhancemen…

Speech Enhancement