paper-with-me

홈 › Papers

RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech Insertion

2022-06-28 · Dacheng Yin, Chuanxin Tang, Yanqing Liu, Xiaoqiang Wang, Zhiyuan Zhao, Yucheng Zhao, Zhiwei Xiong, Sheng Zhao, Chong Luo

This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrary-length speech insertion and even full sentence generation. In the proposed paradigm, global and local factors in speech are explicitly decomposed and separately manipulated to achieve high speaker similarity and continuous prosody. Specifically, we proposed to represent the global factors by multiple tokens, which are extracted by cross-attention operation and then injected back by link-attention operation. Due to the rich representation of global factors, we manage to achieve high speaker similarity in a zero-shot manner. In addition, we introduce a prosody smoothing task to make the local prosody factor context-aware and therefore achieve satisfactory prosody continuity. We further achieve high voice quality with an adversarial training stage. In the subjective test, our method achieves state-of-the-art performance in both naturalness and similarity. Audio samples can be found at https://ydcustc.github.io/retrieverTTS-demo/.

📄 PDF Abstract BibTeX arXiv:2206.13865

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

FastPitchFormant: Source-filter based Decomposed Modeling for Speech Synthesis

2021-06-29 · Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim 외

Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized spee…

Speech Synthesistext-to-speechText to Speech

Fine-grained Noise Control for Multispeaker Speech Synthesis

2022-04-11 · Karolos Nikitaras, Georgios Vamvoukakis, Nikolaos Ellinas, Konstantinos Klapsas 외

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in ord…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech

Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement

2020-11-08 · Daxin Tan, Tan Lee

This paper presents a novel design of neural network system for fine-grained style modeling, transfer and prediction in expressive text-to-speech (TTS) synthesis. Fine-grained modeling is realized by extracting style emb…

DisentanglementSpeech SynthesisStyle Transfertext-to-speech+2

Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective

2024-12-22 · Hankun Wang, Haoran Wang, Yiwei Guo, Zhihan Li 외

Although text-based large language models exhibit human-level writing ability and remarkable intelligence, speech language models (SLMs) still struggle to generate semantically coherent outputs. There are several potenti…

text-to-speechText to Speech

Factor Decomposed Generative Adversarial Networks for Text-to-Image Synthesis

2023-03-24 · Jiguo Li, Xiaobin Liu, Lirong Zheng

Prior works about text-to-image synthesis typically concatenated the sentence embedding with the noise vector, while the sentence embedding and the noise vector are two different factors, which control the different aspe…

Image GenerationSentenceSentence EmbeddingSentence-Embedding