paper-with-me

Papers

Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech

2023-02-27 · Dong Yang, Tomoki Koriyama, Yuki Saito, Takaaki Saeki, Detai Xin, Hiroshi Saruwatari

Pause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers' different styles of inserting silent pauses, which can degrade the performance of the model trained on a multi-speaker speech corpus. To this end, we propose more powerful pause insertion frameworks based on a pre-trained language model. Our approach uses bidirectional encoder representations from transformers (BERT) pre-trained on a large-scale text corpus, injecting speaker embedding to capture various speaker characteristics. We also leverage duration-aware pause insertion for more natural multi-speaker TTS. We develop and evaluate two types of models. The first improves conventional phrasing models on the position prediction of respiratory pauses (RPs), i.e., silent pauses at word transitions without punctuation. It performs speaker-conditioned RP prediction considering contextual information and is used to demonstrate the effect of speaker information on the prediction. The second model is further designed for phoneme-based TTS models and performs duration-aware pause insertion, predicting both RPs and punctuation-indicated pauses (PIPs) that are categorized by duration. The evaluation results show that our models improve the precision and recall of pause insertion and the rhythm of synthetic speech.

📄 PDF Abstract BibTeX arXiv:2302.13652

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingPredictionRhythmtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Synthesizing Dysarthric Speech Using Multi-talker TTS for Dysarthric Speech Recognition

2022-01-27 · Mohammad Soleymanpour, Michael T. Johnson, Rahim Soleymanpour, Jeffrey Berry

Dysarthria is a motor speech disorder often characterized by reduced speech intelligibility through slow, uncoordinated control of speech production muscles. Automatic Speech recognition (ASR) systems may help dysarthric…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3

Isochrony-Controlled Speech-to-Text Translation: A study on translating from Sino-Tibetan to Indo-European Languages

2024-11-11 · Midia Yousefi, Yao Qian, Junkun Chen, Gang Wang 외

End-to-end speech translation (ST), which translates source language speech directly into target language text, has garnered significant attention in recent years. Many ST applications require strict length control to en…

DecoderMachine TranslationSpeech-to-TextSpeech-to-Text Translation+1

Using Pause Information for More Accurate Entity Recognition

2021-09-27 · EMNLP (NLP4ConvAI) 2021 11 · Sahas Dendukuri, Pooja Chitkara, Joel Ruben Antony Moniz, Xiao Yang 외

Entity tags in human-machine dialog are integral to natural language understanding (NLU) tasks in conversational assistants. However, current systems struggle to accurately parse spoken queries with the typical use of te…

Natural Language Understanding

Isochrony-Aware Neural Machine Translation for Automatic Dubbing

2021-12-16 · Derek Tam, Surafel M. Lakew, Yogesh Virkar, Prashant Mathur 외

We introduce the task of isochrony-aware machine translation which aims at generating translations suitable for dubbing. Dubbing of a spoken sentence requires transferring the content as well as the speech-pause structur…

Machine TranslationSentenceTranslation

Naturalization of Text by the Insertion of Pauses and Filler Words

2020-11-07 · Richa Sharma, Parth Vipul Shah, Ashwini M. Joshi

In this article, we introduce a set of methods to naturalize text based on natural human speech. Voice-based interactions provide a natural way of interfacing with electronic systems and are seeing a widespread adaptatio…

Sentencetext-to-speechText to Speech