paper-with-me

홈 › Papers

VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration

2025-01-01 · Stanislav Kirdey

We present VoiceRestore, a novel approach to restoring the quality of speech recordings using flow-matching Transformers trained in a self-supervised manner on synthetic data. Our method tackles a wide range of degradations frequently found in both short and long-form speech recordings, including background noise, reverberation, compression artifacts, and bandwidth limitations - all within a single, unified model. Leveraging conditional flow matching and classifier free guidance, the model learns to map degraded speech to high quality recordings without requiring paired clean and degraded datasets. We describe the training process, the conditional flow matching framework, and the model's architecture. We also demonstrate the model's generalization to real-world speech restoration tasks, including both short utterances and extended monologues or dialogues. Qualitative and quantitative evaluations show that our approach provides a flexible and effective solution for enhancing the quality of speech recordings across varying lengths and degradation types.

📄 PDF Abstract BibTeX arXiv:2501.00794

Code (1)

skirdey/voicerestore 공식 구현 jax

Similar Papers 제목 키워드 기반

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

2025-06-11 · Neta Glazer, Aviv Navon, Yael Segal, Aviv Shamsian 외

Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS mode…

Speech Synthesistext-to-speechText to Speech

Construction of a Large-scale Japanese ASR Corpus on TV Recordings

2021-03-26 · Shintaro Ando, Hiromasa Fujihara

This paper presents a new large-scale Japanese speech corpus for training automatic speech recognition (ASR) systems. This corpus contains over 2,000 hours of speech with transcripts built on Japanese TV recordings and t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models

2025-05-28 · Susan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn 외

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they s…

Speech Synthesis

DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers

2026-03-23 · Tianyu Cao, Helin Wang, Ari Frummer, Yuval Sieradzki 외 arxiv

Recent advances in generative models, such as diffusion and flow matching, have shown strong performance in audio tasks. However, speech enhancement (SE) models are typically trained on limited datasets and evaluated und…

Speech Enhancement

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching

2025-06-19 · Shoutrik Das, Nishant Singh, Arjun Gangwar, S Umesh

Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-r…

Self-Supervised Learning