paper-with-me

Papers

FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial Networks

2025-03-17 · Tong Lei, Qinwen Hu, Ziyao Lin, Andong Li, Rilin Chen, Meng Yu, Dong Yu, Jing Lu

The prevailing method for neural speech enhancement predominantly utilizes fully-supervised deep learning with simulated pairs of far-field noisy-reverberant speech and clean speech. Nonetheless, these models frequently demonstrate restricted generalizability to mixtures recorded in real-world conditions. To address this issue, this study investigates training enhancement models directly on real mixtures. Specifically, we revisit the single-channel far-field to near-field speech enhancement (FNSE) task, focusing on real-world data characterized by low signal-to-noise ratio (SNR), high reverberation, and mid-to-high frequency attenuation. We propose FNSE-SBGAN, a framework that integrates a Schrodinger Bridge (SB)-based diffusion model with generative adversarial networks (GANs). Our approach achieves state-of-the-art performance across various metrics and subjective evaluations, significantly reducing the character error rate (CER) by up to 14.58% compared to far-field signals. Experimental results demonstrate that FNSE-SBGAN preserves superior subjective quality and establishes a new benchmark for real-world far-field speech enhancement. Additionally, we introduce an evaluation framework leveraging matrix rank analysis in the time-frequency domain, providing systematic insights into model performance and revealing the strengths and weaknesses of different generative methods.

📄 PDF Abstract BibTeX arXiv:2503.12936

Code (1)

Taltt/FNSE-SBGAN 공식 구현

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement

2025-08-28 · Mattias Cross, Anton Ragni arxiv

Current flow-based generative speech enhancement methods learn curved probability paths which model a mapping between clean and noisy speech. Despite impressive performance, the implications of curved probability paths a…

Speech Enhancement

Mean Field Optimization Problem Regularized by Fisher Information

2023-02-12 · Julien Claisse, Giovanni Conforti, Zhenjie Ren, SongBo Wang

Recently there is a rising interest in the research of mean field optimization, in particular because of its role in analyzing the training of neural networks. In this paper by adding the Fisher Information as the regula…

ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement

2024-07-28 · Zhong-Qiu Wang

The current dominant approach for neural speech enhancement is via purely-supervised deep learning on simulated pairs of far-field noisy-reverberant speech (i.e., mixtures) and clean speech. The trained models, however, …

Pseudo LabelSpeech Enhancement

Modelling Quantum Channels Carrying Classical Information

2021-12-05 · Indrakshi Dey, Simon L. Cotton

We use the concept of coupled quantum harmonic oscillators to model the propagation environment in which a quantum link carrying either classical or quantum information operates. Using the analogy between the paraxial op…

Dual-Stage Low-Complexity Reconfigurable Speech Enhancement

2021-05-17 · Jun Yang, Nico Brailovsky

This paper proposes a dual-stage, low complexity, and reconfigurable technique to enhance the speech contaminated by various types of noise sources. Driven by input data and audio contents, the proposed dual-stage speech…

Speech Enhancement