paper-with-me

Papers

VoiceBridge: General Speech Restoration with One-step Latent Bridge Models

2025-09-28 · Chi Zhang, Kaiwen Zheng, Zehua Chen, Jun Zhu arxiv

Bridge models have been investigated in speech enhancement but are mostly single-task, with constrained general speech restoration (GSR) capability. In this work, we propose VoiceBridge, a one-step latent bridge model (LBM) for GSR, capable of efficiently reconstructing 48 kHz fullband speech from diverse distortions. To inherit the advantages of data-domain bridge models, we design an energy-preserving variational autoencoder, enhancing the waveform-latent space alignment over varying energy levels. By compressing waveform into continuous latent representations, VoiceBridge models~\textit{various} GSR tasks with a~\textit{single} latent-to-latent generative process backed by a scalable transformer. To alleviate the challenge of reconstructing the high-quality target from distinctively different low-quality priors, we propose a joint neural prior for GSR, uniformly reducing the burden of the LBM in diverse tasks. Building upon these designs, we further investigate bridge training objective by jointly tuning LBM, decoder and discriminator together, transforming the model from a denoiser to generator and enabling \textit{one-step GSR without distillation}. Extensive validation across in-domain (\textit{e.g.}, denoising and super-resolution) and out-of-domain tasks (\textit{e.g.}, refining synthesized speech) and datasets demonstrates the superior performance of VoiceBridge. Demos: https://VoiceBridgedemo.github.io/.

📄 PDF Abstract BibTeX arXiv:2509.25275

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

DM: Dual-path Magnitude Network for General Speech Restoration

2024-09-13 · Da-Hee Yang, Dail Kim, Joon-Hyuk Chang, Jeonghwan Choi 외

In this paper, we introduce a novel general speech restoration model: the Dual-path Magnitude (DM) network, designed to address multiple distortions including noise, reverberation, and bandwidth degradation effectively. …

Decoder

High-Resolution Speech Restoration with Latent Diffusion Model

2024-09-17 · Tushar Dhyani, Florian Lux, Michele Mancusi, Giorgio Fabbro 외

Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstructi…

modelSpeech Enhancement

FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech Restoration

2025-05-03 · Da-Hee Yang, Jaeuk Lee, Joon-Hyuk Chang

We introduce FLOWER, a novel conditioning method designed for speech restoration that integrates Gaussian guidance into generative frameworks. By transforming clean speech into a predefined prior distribution (e.g., Gaus…

Automatic Restoration of Diacritics for Speech Data Sets

2023-11-15 · Sara Shatnawi, Sawsan Alqahtani, Hanan Aldarmaki

Automatic text-based diacritic restoration models generally have high diacritic error rates when applied to speech transcripts as a result of domain and style shifts in spoken language. In this work, we explore the possi…

LCUDiff: Latent Capacity Upgrade Diffusion for Faithful Human Body Restoration

2026-02-04 · Jue Gong, Zihan Zhou, Jingkai Wang, Shu Li 외 arxiv

Existing methods for restoring degraded human-centric images often struggle with insufficient fidelity, particularly in human body restoration (HBR). Recent diffusion-based restoration methods commonly adapt pre-trained …