paper-with-me

홈 › Papers

Robust One-step Speech Enhancement via Consistency Distillation

2025-07-08 · Liang Xu, Longfei Felix Yan, W. Bastiaan Kleijn

Diffusion models have shown strong performance in speech enhancement, but their real-time applicability has been limited by multi-step iterative sampling. Consistency distillation has recently emerged as a promising alternative by distilling a one-step consistency model from a multi-step diffusion-based teacher model. However, distilled consistency models are inherently biased towards the sampling trajectory of the teacher model, making them less robust to noise and prone to inheriting inaccuracies from the teacher model. To address this limitation, we propose ROSE-CD: Robust One-step Speech Enhancement via Consistency Distillation, a novel approach for distilling a one-step consistency model. Specifically, we introduce a randomized learning trajectory to improve the model's robustness to noise. Furthermore, we jointly optimize the one-step model with two time-domain auxiliary losses, enabling it to recover from teacher-induced errors and surpass the teacher model in overall performance. This is the first pure one-step consistency distillation model for diffusion-based speech enhancement, achieving 54 times faster inference speed and superior performance compared to its 30-step teacher model. Experiments on the VoiceBank-DEMAND dataset demonstrate that the proposed model achieves state-of-the-art performance in terms of speech quality. Moreover, its generalization ability is validated on both an out-of-domain dataset and real-world noisy recordings.

📄 PDF Abstract BibTeX arXiv:2507.05688

Code (1)

LiangXu123/Robust-One-step-Speech-Enhancement-via-Consistency-Distillation-ROSE-CD- 공식 구현

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Consistency Models 설명 없음

Similar Papers 제목 키워드 기반

Two-Step Knowledge Distillation for Tiny Speech Enhancement

2023-09-15 · Rayan Daod Nathoo, Mikolaj Kegler, Marko Stamenovic

Tiny, causal models are crucial for embedded audio machine learning applications. Model compression can be achieved via distilling knowledge from a large teacher into a smaller student model. In this work, we propose a n…

Knowledge DistillationModel CompressionSpeech Enhancement

MeanFlowSE: one-step generative speech enhancement via conditional mean flow

2025-09-18 · Duojia Li, Shenghui Lu, Hongchen Pan, Zongyi Zhan 외 arxiv

Multistep inference is a bottleneck for real-time generative speech enhancement because flow- and diffusion-based systems learn an instantaneous velocity field and therefore rely on iterative ordinary differential equati…

Knowledge DistillationSpeech Enhancement

Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation

2023-05-24 · Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

Audio-visual speech enhancement (AV-SE) aims to enhance degraded speech along with extra visual information such as lip videos, and has been shown to be more effective than audio-only speech enhancement. This paper propo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge DistillationSpeech Enhancement+2

Compose Yourself: Average-Velocity Flow Matching for One-Step Speech Enhancement

2025-09-19 · Gang Yang, Yue Lei, Wenxin Tai, Jin Wu 외 arxiv

Diffusion and flow matching (FM) models have achieved remarkable progress in speech enhancement (SE), yet their dependence on multi-step generation is computationally expensive and vulnerable to discretization errors. Re…

Speech Enhancement

Multi-View Attention Transfer for Efficient Speech Enhancement

2022-08-22 · WooSeok Shin, Hyun Joon Park, Jin Sob Kim, Byung Hoon Lee 외

Recent deep learning models have achieved high performance in speech enhancement; however, it is still challenging to obtain a fast and low-complexity model without significant performance degradation. Previous knowledge…

Knowledge DistillationSpeech Enhancement