paper-with-me

홈 › Papers

Miipher: A Robust Speech Restoration Model Integrating Self-Supervised Speech and Text Representations

2023-03-03 · Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding, Kohei Yatabe, Nobuyuki Morioka, Yu Zhang, Wei Han, Ankur Bapna, Michiel Bacchiani

Speech restoration (SR) is a task of converting degraded speech signals into high-quality ones. In this study, we propose a robust SR model called Miipher, and apply Miipher to a new SR application: increasing the amount of high-quality training data for speech generation by converting speech samples collected from the Web to studio-quality. To make our SR model robust against various degradation, we use (i) a speech representation extracted from w2v-BERT for the input feature, and (ii) a text representation extracted from transcripts via PnG-BERT as a linguistic conditioning feature. Experiments show that Miipher (i) is robust against various audio degradation and (ii) enable us to train a high-quality text-to-speech (TTS) model from restored speech samples collected from the Web. Audio samples are available at our demo page: google.github.io/df-conformer/miipher/

📄 PDF Abstract BibTeX arXiv:2303.01664

Code (1)

Wataru-Nakata/miipher pytorch

Tasks

Speech DenoisingSpeech Enhancementtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration

2025-05-07 · Shigeki Karita, Yuma Koizumi, Heiga Zen, Haruko Ishikawa 외

Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, for training data cleaning for large-scal…

Computational Efficiency

FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks

2024-08-12 · Min Ma, Yuma Koizumi, Shigeki Karita, Heiga Zen 외

This paper introduces FLEURS-R, a speech restoration applied version of the Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) corpus. FLEURS-R maintains an N-way parallel speech corpus in 102 l…

Few-Shot Learningtext-to-speechText to Speech

An empirical study on speech restoration guided by self supervised speech representation

2023-05-30 · Jaeuk Byun, Youna Ji, Soo Whan Chung, Soyeon Choe 외

Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adverse…

Representation LearningSpeech Representation Learning

MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets

2022-11-14 · Ziyang Ma, Zhisheng Zheng, Changli Tang, Yujin Wang 외

In this paper, we provide a new perspective on self-supervised speech models from how the training targets are obtained. We generalize the targets extractor into Offline Targets Extractor (Off-TE) and Online Targets Extr…

Automatic Speech RecognitionMulti-Task LearningRepresentation LearningSelf-Supervised Learning+2

Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches

2025-08-11 · Ahmed Aboeitta, Ahmed Sharshar, Youssef Nafea, Shady Shehata arxiv

Speech Recognition (ASR) due to phoneme distortions and high variability. While self-supervised ASR models like Wav2Vec, HuBERT, and Whisper have shown promise, their effectiveness in dysarthric speech remains unclear. T…

Speech Recognition