paper-with-me

홈 › Papers

On the Interplay Between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis

2021-10-04 · Cheng-I Jeff Lai, Erica Cooper, Yang Zhang, Shiyu Chang, Kaizhi Qian, Yi-Lun Liao, Yung-Sung Chuang, Alexander H. Liu, Junichi Yamagishi, David Cox, James Glass

Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction networks and vocoders. We thoroughly investigate the tradeoffs between sparsity and its subsequent effects on synthetic speech. Additionally, we explored several aspects of TTS pruning: amount of finetuning data versus sparsity, TTS-Augmentation to utilize unspoken text, and combining knowledge distillation and pruning. Our findings suggest that not only are end-to-end TTS models highly prunable, but also, perhaps surprisingly, pruned TTS models can produce synthetic speech with equal or higher naturalness and intelligibility, with similar prosody. All of our experiments are conducted on publicly available models, and findings in this work are backed by large-scale subjective tests and objective measures. Code and 200 pruned models are made available to facilitate future research on efficiency in TTS.

📄 PDF Abstract BibTeX arXiv:2110.01147

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationSpeech Synthesistext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Pruning 설명 없음
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units

2022-11-12 · Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky

We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody and language content, leading to the deg…

RhythmVoice Conversion

Lightweight and perceptually-guided voice conversion for electro-laryngeal speech

2026-01-07 · Benedikt Mayrhofer, Franz Pernkopf, Philipp Aichinger, Martin Hagmüller arxiv

Electro-laryngeal (EL) speech is characterized by constant pitch, limited prosody, and mechanical noise, reducing naturalness and intelligibility. We propose a lightweight adaptation of the state-of-the-art StreamVC fram…

Voice Conversion

Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech

2022-05-09 · ACL 2022 5 · Yang Li, Cheng Yu, Guangzhi Sun, Hua Jiang 외

Modelling prosody variation is critical for synthesizing natural and expressive speech in end-to-end text-to-speech (TTS) systems. In this paper, a cross-utterance conditional VAE (CUC-VAE) is proposed to estimate a post…

Diversitytext-to-speechText to Speech

Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Modelling prosody variation is critical for synthesizing natural and expressive speech in end-to-end text-to-speech (TTS) systems. In this paper, a cross-utterance conditional VAE (CUC-VAE) is proposed to estimate a post…

Diversitytext-to-speechText to Speech

Automatic Prosody Prediction for Chinese Speech Synthesis using BLSTM-RNN and Embedding Features

2015-11-02 · Chuang Ding, Lei Xie, Jie Yan, Weini Zhang 외

Prosody affects the naturalness and intelligibility of speech. However, automatic prosody prediction from text for Chinese speech synthesis is still a great challenge and the traditional conditional random fields (CRF) b…

Feature EngineeringProsody PredictionSpeech Synthesis