paper-with-me

Papers

Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement

2026-06-19 · Nasser-Eddine Monir, Paul Magron, Romain Serizel arxiv

Conventional training losses for speech enhancement based on the signal-to-distortion ratio (SDR) treat all time-frequency (TF) regions uniformly, overlooking the fine-grained spectral cues that are relevant to specific phoneme intelligibility. We propose a TF weighting framework that modulates the SDR objective based on local speech presence, speech-to-interference ratio (SIR), and spectral flux. By integrating these factors into a differentiable objective, the framework emphasizes TF bins with high speech-noise competition while also accounting for transient cues such as consonant bursts. Experimental results show that our approach improves objective frequency-weighted enhancement metrics, as well as phoneme recognition accuracy, particularly for consonants. Spectral analysis shows better reconstruction of mid-frequency structures at less adverse SIR.

📄 PDF Abstract BibTeX arXiv:2606.21635

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement

2025-06-23 · Nasser-Eddine Monir, Paul Magron, Romain Serizel

Recent advances in deep learning have significantly improved multichannel speech enhancement algorithms, yet conventional training loss functions such as the scale-invariant signal-to-distortion ratio (SDR) may fail to p…

Speech Enhancement

FRICATIVE PHONEME DETECTION WITH ZERO DELAY

2019-09-25 · Metehan Yurt, Alberto N. Escalante B., Veniamin I. Morgenshtern

People with high-frequency hearing loss rely on hearing aids that employ frequency lowering algorithms. These algorithms shift some of the sounds from the high frequency band to the lower frequency band where the sounds …

Improved Balanced Classification with Theoretically Grounded Loss Functions

2025-12-30 · Corinna Cortes, Mehryar Mohri, Yutao Zhong arxiv

The balanced loss is a widely adopted objective for multi-class classification under class imbalance. By assigning equal importance to all classes, regardless of their frequency, it promotes fairness and ensures that min…

Multi-class Classification

Phoneme-based Distribution Regularization for Speech Enhancement

2021-04-08 · Yajing Liu, Xiulian Peng, Zhiwei Xiong, Yan Lu

Existing speech enhancement methods mainly separate speech from noises at the signal level or in the time-frequency domain. They seldom pay attention to the semantic information of a corrupted signal. In this paper, we a…

Speech Enhancement

Temporal Dynamic Convolutional Neural Network for Text-Independent Speaker Verification and Phonemetic Analysis

2021-10-07 · Seong-Hu Kim, Hyeonuk Nam, Yong-Hwa Park

In the field of text-independent speaker recognition, dynamic models that adapt along the time axis have been proposed to consider the phoneme-varying characteristics of speech. However, a detailed analysis of how dynami…

Speaker RecognitionSpeaker VerificationText-Independent Speaker RecognitionText-Independent Speaker Verification