paper-with-me

홈 › Papers

Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features

2019-11-21 · Siddharth Gururani, Kilol Gupta, Dhaval Shah, Zahra Shakeri, Jervis Pinto

This paper presents a simple yet effective method to achieve prosody transfer from a reference speech signal to synthesized speech. The main idea is to incorporate well-known acoustic correlates of prosody such as pitch and loudness contours of the reference speech into a modern neural text-to-speech (TTS) synthesizer such as Tacotron2 (TC2). More specifically, a small set of acoustic features are extracted from reference audio and then used to condition a TC2 synthesizer. The trained model is evaluated using subjective listening tests and a novel objective evaluation of prosody transfer is proposed. Listening tests show that the synthesized speech is rated as highly natural and that prosody is successfully transferred from the reference speech signal to the synthesized signal.

📄 PDF Abstract BibTeX arXiv:1911.09645

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Global Rhythm Style Transfer Without Text Transcriptions

2021-06-16 · Kaizhi Qian, Yang Zhang, Shiyu Chang, JinJun Xiong 외

Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody information. Two major components of pro…

Representation LearningRhythmStyle Transfer

Pitchtron: Towards audiobook generation from ordinary people's voices

2020-05-21 · Interspeech 2020 5 · Sunghee Jung, Hoirin Kim

In this paper, we explore prosody transfer for audiobook generation under rather realistic condition where training DB is plain audio mostly from multiple ordinary people and reference audio given during inference is fro…

Decoder

Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2

2026-03-12 · Suvendu Sekhar Mohanty arxiv

We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual tra…

Speech Synthesis

A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units

2022-11-12 · Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky

We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody and language content, leading to the deg…

RhythmVoice Conversion

Robust and fine-grained prosody control of end-to-end speech synthesis

2018-11-06 · Young-Gun Lee, Taesu Kim

We propose prosody embeddings for emotional and expressive speech synthesis networks. The proposed methods introduce temporal structures in the embedding networks, thus enabling fine-grained control of the speaking style…

Expressive Speech SynthesisSpeech Synthesis