paper-with-me

홈 › Papers

Pitchtron: Towards audiobook generation from ordinary people's voices

2020-05-21 · Interspeech 2020 5 · Sunghee Jung, Hoirin Kim

In this paper, we explore prosody transfer for audiobook generation under rather realistic condition where training DB is plain audio mostly from multiple ordinary people and reference audio given during inference is from professional and richer in prosody than training DB. To be specific, we explore transferring Korean dialects and emotive speech even though training set is mostly composed of standard and neutral Korean. We found that under this setting, original global style token method generates undesirable glitches in pitch, energy and pause length. To deal with this issue, we propose two models, hard and soft pitchtron and release the toolkit and corpus that we have developed. Hard pitchtron uses pitch as input to the decoder while soft pitchtron uses pitch as input to the prosody encoder. We verify the effectiveness of proposed models with objective and subjective tests. AXY score over GST is 2.01 and 1.14 for hard pitchtron and soft pitchtron respectively.

📄 PDF Abstract BibTeX arXiv:2005.10456

Code (1)

hash2430/pitchtron 공식 구현 pytorch

Tasks

Decoder

Similar Papers 제목 키워드 기반

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

2025-05-19 · Kyeongman Park, Seongho Joo, Kyomin Jung

We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook …

Sentence

Using Audio Books for Training a Text-to-Speech System

2014-05-01 · LREC 2014 5 · Chalam, Aimilios aris, Pirros Tsiakoulis, Sotiris Karabetsos 외

Creating new voices for a TTS system often requires a costly procedure of designing and recording an audio corpus, a time consuming and effort intensive task. Using publicly available audiobooks as the raw material of a …

DiversitySpeech Synthesistext-to-speechText to Speech

Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices

2022-06-01 · LREC 2022 6 · Liisi Piits, Hille Pajupuu, Heete Sahkai, Rene Altrov 외

Synthetic voices are increasingly used in applications that require a conversational speaking style, raising the question as to which type of training data yields the most suitable speaking style for such applications. T…

Sentencetext-to-speechText to Speech

TTS Skins: Speaker Conversion via ASR

2019-04-18 · Adam Polyak, Lior Wolf, Yaniv Taigman

We present a fully convolutional wav-to-wav network for converting between speakers' voices, without relying on text. Our network is based on an encoder-decoder architecture, where the encoder is pre-trained for the task…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2