paper-with-me

홈 › Papers

Aligning Generative Speech Enhancement with Perceptual Feedback

2025-07-14 · Haoyang Li, Nana Hou, Yuchen Hu, Jixun Yao, Sabato Marco Siniscalchi, Xuyi Zhuang, Deheng Ye, Wei Yang, Eng Siong Chng arxiv

Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives that weakly reflect human perception. This mismatch limits progress, as optimizing signal accuracy does not always improve naturalness or listening comfort. We address this gap by introducing a perceptually aligned LM-based SE approach. Our method applies Direct Preference Optimization (DPO) with UTMOS, a neural MOS predictor, as a proxy for human ratings, directly steering models toward perceptually preferred outputs. This design directly connects model training to perceptual quality and is broadly applicable within LM-based SE frameworks. On the Deep Noise Suppression Challenge 2020 test sets, our approach consistently improves speech quality metrics, achieving relative gains of up to 56%. To our knowledge, this is the first integration of perceptual feedback into LM-based SE and the first application of DPO in the SE domain, establishing a new paradigm for perceptually aligned enhancement with SE.

📄 PDF Abstract BibTeX arXiv:2507.09929

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Diffiner: A Versatile Diffusion-based Generative Refiner for Speech Enhancement

2022-10-27 · Ryosuke Sawata, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka 외

Although deep neural network (DNN)-based speech enhancement (SE) methods outperform the previous non-DNN-based ones, they often degrade the perceptual quality of generated outputs. To tackle this problem, we introduce a …

DenoisingSpeech Enhancement

Investigating Training Objectives for Generative Speech Enhancement

2024-09-16 · Julius Richter, Danilo de Oliveira, Timo Gerkmann

Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learni…

Speech Enhancement

FINALLY: fast and universal speech enhancement with studio-like quality

2024-10-08 · Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov 외

In this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, and microphone artifacts. We revisit the u…

Speech Enhancement

Predictive-Generative Drift Decomposition for Speech Enhancement and Separation

2026-05-07 · Julius Richter, Yoshiki Masuyama, Christoph Boeddeker, Takahiro Edo 외 arxiv

We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpolant Prior for Speech (SIPS), builds on s…

Speech EnhancementSpeech Separation

DCNGAN: A Deformable Convolutional-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed Video

2022-01-22 · Saiping Zhang, Luis Herranz, Marta Mrak, Marc Gorriz Blanch 외

In this paper, we propose a deformable convolution-based generative adversarial network (DCNGAN) for perceptual quality enhancement of compressed videos. DCNGAN is also adaptive to the quantization parameters (QPs). Comp…

Generative Adversarial NetworkQuantization