paper-with-me

홈 › Papers

Using Audio Books for Training a Text-to-Speech System

2014-05-01 · LREC 2014 5 · Chalam, Aimilios aris, Pirros Tsiakoulis, Sotiris Karabetsos, Spyros Raptis

Creating new voices for a TTS system often requires a costly procedure of designing and recording an audio corpus, a time consuming and effort intensive task. Using publicly available audiobooks as the raw material of a spoken corpus for such systems creates new perspectives regarding the possibility of creating new synthetic voices quickly and with limited effort. This paper addresses the issue of creating new synthetic voices based on audiobook data in an automated method. As an audiobook includes several types of speech, such as narration, character playing etc., special care is given in identifying the data subset that leads to a more neutral and general purpose synthetic voice. The main goal is to identify and address the effect the audiobook speech diversity on the resulting TTS system. Along with the methodology for coping with this diversity in the speech data, we also describe a set of experiments performed in order to investigate the efficiency of different approaches for automatic data pruning. Further plans for exploiting the diversity of the speech incorporated in an audiobook are also described in the final section and conclusions are drawn.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DiversitySpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2

Large-Scale Automatic Audiobook Creation

2023-09-07 · Brendan Walsh, Mark Hamilton, Greg Newby, Xi Wang 외

An audiobook can dramatically improve a work of literature's accessibility and improve reader engagement. However, audiobooks can take hundreds of hours of human effort to create, edit, and publish. In this work, we pres…

text-to-speechText to Speech

IESTAC: English-Italian Parallel Corpus for End-to-End Speech-to-Text Machine Translation

2020-11-01 · EMNLP (nlpbt) 2020 11 · Giuseppe Della Corte, Sara Stymne

We discuss a set of methods for the creation of IESTAC: a English-Italian speech and text parallel corpus designed for the training of end-to-end speech-to-text machine translation models and publicly released as part of…

Dynamic Time WarpingMachine TranslationSentenceSentence Embeddings+4

Towards Fully Automatic Annotation of Audio Books for TTS

2012-05-01 · LREC 2012 5 · Olivier Boeffard, Laure Charonnat, S{\'e}bastien Le Maguer, Damien Lolive

Building speech corpora is a first and crucial step for every text-to-speech synthesis system. Nowadays, the use of statistical models implies the use of huge sized corpora that need to be recorded, transcribed, annotate…

Speech RecognitionSpeech Synthesistext-to-speechText to Speech+1

End-to-End Automatic Speech Translation of Audiobooks

2018-02-12 · Alexandre Bérard, Laurent Besacier, Ali Can Kocabiyikoglu, Olivier Pietquin

We investigate end-to-end speech-to-text translation on a corpus of audiobooks specifically augmented for this task. Previous works investigated the extreme case where source language transcription is not available durin…

automatic-speech-translationSpeech-to-TextSpeech-to-Text TranslationTranslation