paper-with-me

Papers

The Potential of Neural Speech Synthesis-based Data Augmentation for Personalized Speech Enhancement

2022-11-14 · Anastasia Kuznetsova, Aswin Sivaraman, Minje Kim

With the advances in deep learning, speech enhancement systems benefited from large neural network architectures and achieved state-of-the-art quality. However, speaker-agnostic methods are not always desirable, both in terms of quality and their complexity, when they are to be used in a resource-constrained environment. One promising way is personalized speech enhancement (PSE), which is a smaller and easier speech enhancement problem for small models to solve, because it focuses on a particular test-time user. To achieve the personalization goal, while dealing with the typical lack of personal data, we investigate the effect of data augmentation based on neural speech synthesis (NSS). In the proposed method, we show that the quality of the NSS system's synthetic data matters, and if they are good enough the augmented dataset can be used to improve the PSE system that outperforms the speaker-agnostic baseline. The proposed PSE systems show significant complexity reduction while preserving the enhancement quality.

📄 PDF Abstract BibTeX arXiv:2211.07493

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationSpeech EnhancementSpeech Synthesis

Similar Papers 제목 키워드 기반

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

2026-03-04 · Youngwon Choi, Jinwoo Oh, Hwayeon Kim, Hyeonyu Kim arxiv

We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically dive…

Data AugmentationSpeech Synthesis

Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement

2025-01-23 · Jae-Sung Bae, Anastasia Kuznetsova, Dinesh Manocha, John Hershey 외

This paper presents a new challenge that calls for zero-shot text-to-speech (TTS) systems to augment speech data for the downstream task, personalized speech enhancement (PSE), as part of the Generative Data Augmentation…

Data AugmentationSpeech EnhancementSpeech SynthesisSynthetic Data Generation+2

Empirical Study Incorporating Linguistic Knowledge on Filled Pauses for Personalized Spontaneous Speech Synthesis

2022-10-14 · Yuta Matsunaga, Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari

We present a comprehensive empirical study for personalized spontaneous speech synthesis on the basis of linguistic knowledge. With the advent of voice cloning for reading-style speech synthesis, a new voice cloning para…

Speech SynthesisVoice Cloning

Residual-guided Personalized Speech Synthesis based on Face Image

2022-04-01 · Jianrong Wang, Zixuan Wang, Xiaosheng Hu, XueWei Li 외

Previous works derive personalized speech features by training the model on a large dataset composed of his/her audio sounds. It was reported that face information has a strong link with the speech sound. Thus in this wo…

Speech Synthesis

Speech Synthesis as Augmentation for Low-Resource ASR

2020-12-23 · Deblin Bagchi, Shannon Wotherspoon, Zhuolin Jiang, Prasanna Muthukumar

Speech synthesis might hold the key to low-resource speech recognition. Data augmentation techniques have become an essential part of modern speech recognition training. Yet, they are simple, naive, and rarely reflect re…

Data Augmentationspeech-recognitionSpeech RecognitionSpeech Synthesis