paper-with-me

Papers

Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement

2025-06-23 · Nasser-Eddine Monir, Paul Magron, Romain Serizel

Recent advances in deep learning have significantly improved multichannel speech enhancement algorithms, yet conventional training loss functions such as the scale-invariant signal-to-distortion ratio (SDR) may fail to preserve fine-grained spectral cues essential for phoneme intelligibility. In this work, we propose perceptually-informed variants of the SDR loss, formulated in the time-frequency domain and modulated by frequency-dependent weighting schemes. These weights are designed to emphasize time-frequency regions where speech is prominent or where the interfering noise is particularly strong. We investigate both fixed and adaptive strategies, including ANSI band-importance weights, spectral magnitude-based weighting, and dynamic weighting based on the relative amount of speech and noise. We train the FaSNet multichannel speech enhancement model using these various losses. Experimental results show that while standard metrics such as the SDR are only marginally improved, their perceptual frequency-weighted counterparts exhibit a more substantial improvement. Besides, spectral and phoneme-level analysis indicates better consonant reconstruction, which points to a better preservation of certain acoustic cues.

📄 PDF Abstract BibTeX arXiv:2506.18714

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement

2026-06-19 · Nasser-Eddine Monir, Paul Magron, Romain Serizel arxiv

Conventional training losses for speech enhancement based on the signal-to-distortion ratio (SDR) treat all time-frequency (TF) regions uniformly, overlooking the fine-grained spectral cues that are relevant to specific …

Speech Enhancement

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages

2026-04-20 · V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard Jäger arxiv

We present a phoneme-level analysis of automatic speech recognition (ASR) for two low-resourced and phonologically complex East Caucasian languages, Archi and Rutul, based on curated and standardized speech-transcript re…

Speech Recognition

The Distribution of Phoneme Frequencies across the World's Languages: Macroscopic and Microscopic Information-Theoretic Models

2026-03-03 · Fermín Moscoso del Prado Martín, Suchir Salhan arxiv

We demonstrate that the frequency distribution of phonemes across languages can be explained at both macroscopic and microscopic levels. Macroscopically, phoneme rank-frequency distributions closely follow the order stat…

Phoneme Level Language Models for Sequence Based Low Resource ASR

2019-02-20 · Siddharth Dalmia, Xinjian Li, Alan W. black, Florian Metze

Building multilingual and crosslingual models help bring different languages together in a language universal space. It allows models to share parameters and transfer knowledge across languages, enabling faster and bette…

Language ModelingLanguage Modelling

Phoneme-based Distribution Regularization for Speech Enhancement

2021-04-08 · Yajing Liu, Xiulian Peng, Zhiwei Xiong, Yan Lu

Existing speech enhancement methods mainly separate speech from noises at the signal level or in the time-frequency domain. They seldom pay attention to the semantic information of a corrupted signal. In this paper, we a…

Speech Enhancement