paper-with-me

Papers

Phoneme-based Distribution Regularization for Speech Enhancement

2021-04-08 · Yajing Liu, Xiulian Peng, Zhiwei Xiong, Yan Lu

Existing speech enhancement methods mainly separate speech from noises at the signal level or in the time-frequency domain. They seldom pay attention to the semantic information of a corrupted signal. In this paper, we aim to bridge this gap by extracting phoneme identities to help speech enhancement. Specifically, we propose a phoneme-based distribution regularization (PbDr) for speech enhancement, which incorporates frame-wise phoneme information into speech enhancement network in a conditional manner. As different phonemes always lead to different feature distributions in frequency, we propose to learn a parameter pair, i.e. scale and bias, through a phoneme classification vector to modulate the speech enhancement network. The modulation parameter pair includes not only frame-wise but also frequency-wise conditions, which effectively map features to phoneme-related distributions. In this way, we explicitly regularize speech enhancement features by recognition vectors. Experiments on public datasets demonstrate that the proposed PbDr module can not only boost the perceptual quality for speech enhancement but also the recognition accuracy of an ASR system on the enhanced speech. This PbDr module could be readily incorporated into other speech enhancement networks as well.

📄 PDF Abstract BibTeX arXiv:2104.03759

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Phoneme-Based Ratio Mask Estimation for Reverberant Speech Enhancement in Cochlear Implant Processors

2021-05-28 · Kevin M. Chu, Leslie M. Collins, Boyla O. Mainsah

Cochlear implant (CI) users have considerable difficulty in understanding speech in reverberant listening environments. Time-frequency (T-F) masking is a common technique that aims to improve speech intelligibility by mu…

SentenceSpeech Enhancement

Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement

2026-06-19 · Nasser-Eddine Monir, Paul Magron, Romain Serizel arxiv

Conventional training losses for speech enhancement based on the signal-to-distortion ratio (SDR) treat all time-frequency (TF) regions uniformly, overlooking the fine-grained spectral cues that are relevant to specific …

Speech Enhancement

PAAPLoss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement

2023-02-16 · Muqiao Yang, Joseph Konan, David Bick, Yunyang Zeng 외

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in …

Speech EnhancementTime SeriesTime Series Analysis

Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement

2025-06-23 · Nasser-Eddine Monir, Paul Magron, Romain Serizel

Recent advances in deep learning have significantly improved multichannel speech enhancement algorithms, yet conventional training loss functions such as the scale-invariant signal-to-distortion ratio (SDR) may fail to p…

Speech Enhancement

Incorporating Broad Phonetic Information for Speech Enhancement

2020-08-13 · Yen-Ju Lu, Chien-Feng Liao, Xugang Lu, Jeih-weih Hung 외

In noisy conditions, knowing speech contents facilitates listeners to more effectively suppress background noise components and to retrieve pure speech signals. Previous studies have also confirmed the benefits of incorp…

DenoisingSpeech Enhancement